unconst/Affine-5czsc2fc98-r1064-vera-odpo-midrank-midbeta-midctx-ultraextra-ep4-hilr-merged
The unconst/Affine-5czsc2fc98-r1064-vera-odpo-midrank-midbeta-midctx-ultraextra-ep4-hilr-merged model is a 35.1 billion parameter Affine SN120 challenger, fine-tuned using offline DPO on Reason-ranked duel pairs. Optimized for the Affine evalsrv Reason v4 duel, it focuses on improving preference for thoughts that raise teacher-side Reason. This model is specifically designed for Affine miner submissions and is not intended as a general-purpose chat model, featuring a 32768 token context length.
Loading preview...
Model Overview
This model, unconst/Affine-5czsc2fc98-r1064-vera-odpo-midrank-midbeta-midctx-ultraextra-ep4-hilr-merged, is a 35.1 billion parameter Affine SN120 challenger, specifically developed for the Reason v4 evaluation. It is based on the vera6/affine-5g4yy75zuz-t6 parent model and was trained using an offline DPO (Direct Preference Optimization) method, rather than Supervised Fine-Tuning (SFT) or online GRPO.
Key Characteristics & Training
- Optimization Target: The model was optimized to enhance preference for thoughts that increase teacher-side Reason scores, utilizing a tempered multi-sample log-mean-exp over k=3 teacher references.
- Training Data: Fine-tuned on 604 lines of Soft Mid Mid Soft → MidCtx filtered duel preference pairs (
dpo_duel_reason.jsonl). - Hyperparameters: Notable hyperparameters include LoRA r=32 (MidRank), α=128 (HiAlpha), β=0.1 (MidBeta), a high learning rate (lr=2e-6, HiLR), and a significant maximum context length of 8192 tokens (MidCtx).
- Performance: Achieved a margin of +0.006632 against the live king
reign36underwvk=7, indicating a successful challenge and licensing as a WIN / Stage-5 model.
Intended Use
This model is explicitly designed for Affine miner submissions and evalsrv Reason v4 duels. It is not intended as a general chat model and its performance is optimized for its specific evaluation context.