ledgernova/sn120-f3314c7cdc17
ledgernova/sn120-f3314c7cdc17 is a 35.1 billion parameter Affine SN120 challenger model developed by ledgernova, specifically optimized for the Reason v4 evaluation server duel. This model was trained using offline DPO on Reason-ranked duel pairs, focusing on preferences for thoughts that enhance teacher-side Reason. With a 32768 token context length, it is designed for specific evaluation tasks rather than general chat applications.
Loading preview...
Overview
ledgernova/sn120-f3314c7cdc17 is a 35.1 billion parameter Affine SN120 challenger model, developed by ledgernova, specifically engineered for the Reason v4 evaluation server duel. It is not intended as a general-purpose chat model but rather for highly specialized evaluation tasks.
Key Training Details
This model was trained using offline DPO (Direct Preference Optimization) on Reason-ranked duel pairs, rather than traditional SFT (Supervised Fine-Tuning) or online GRPO. The optimization focused on enhancing preferences for thoughts that improve teacher-side Reason scores. Key hyperparameters included a LoRA rank of 32, alpha of 128, beta of 0.3, and a learning rate of 1e-6. It was trained for 3 epochs with a maximum context length of 6144 tokens.
Performance & Intended Use
During its development, this checkpoint demonstrated a winning margin of +0.002137 against the live king reign34 model under wvk=7, achieving a B pass of 0.304. Its primary intended use is as an SN120 Affine miner submission and for evalsrv Reason v4 duels, reflecting its highly specialized nature.