ledgernova/sn120-8ef1b06acfe8
The ledgernova/sn120-8ef1b06acfe8 model is a 35.1 billion parameter Affine SN120 challenger, specifically optimized for Reason v4 evaluations. Trained using offline DPO on Reason-ranked duel pairs, it demonstrates a preference for thoughts that enhance teacher-side Reason. This model is designed for specialized mining submissions and evaluation server Reason v4 duels, rather than general chat applications.
Loading preview...
Overview
This model, ledgernova/sn120-8ef1b06acfe8, is an Affine SN120 challenger, a specialized 35.1 billion parameter model derived from the vera6/affine-5g4yy75zuz-t6 base. It was developed using offline DPO (Direct Preference Optimization), focusing on optimizing for Reason v4 evaluations.
Key Training Details
- Methodology: Offline DPO on Reason-ranked duel pairs, not SFT or online GRPO.
- Optimization Goal: To prefer thoughts that increase teacher-side Reason, using a tempered multi-sample log-mean-exp over teacher references.
- Data: Filtered duel preference pairs from
dpo_duel_reason.jsonlwith SoftCtx filtering. - Key Hyperparameters: Notable settings include LoRA r=32 (MidRank), α=128 (HiAlpha), β=0.3 (HiBeta), a very low learning rate (lr=5e-7, UltraLoLR), and a max_len of 12288 (SoftCtx).
- Performance: Achieved a win against the live king reign36 with a margin of +0.004951 and a z-score of 2.399 on n=80 evaluations for Reason v4.
Intended Use
This model is specifically designed for SN120 Affine miner submissions and evaluation server Reason v4 duels. It is not intended as a general chat model.