unconst/Affine-5czsc2fc98-r1032-vera-odpo-midrank-hibeta-shortctx-ultraextra-ep4-midlr-merged
The unconst/Affine-5czsc2fc98-r1032-vera-odpo-midrank-hibeta-shortctx-ultraextra-ep4-midlr-merged model is a 35.1 billion parameter Affine SN120 challenger, specifically fine-tuned using offline DPO for enhanced performance in Reason v4 evaluations. It optimizes for thoughts that raise teacher-side Reason, utilizing a short context length of 6144 tokens and a LoRA rank of 32. This model is primarily intended for SN120 Affine miner submissions and evalsrv Reason v4 duels, rather than general chat applications.
Loading preview...
Overview
This model, unconst/Affine-5czsc2fc98-r1032-vera-odpo-midrank-hibeta-shortctx-ultraextra-ep4-midlr-merged, is a 35.1 billion parameter Affine SN120 challenger. It was developed for the Reason v4 evaluation (weight_version_key=7) and is based on the vera6/affine-5g4yy75zuz-t6 parent model. The training involved an offline DPO (Direct Preference Optimization) method, specifically targeting preference for thoughts that improve teacher-side Reason scores.
Key Characteristics
- Optimization Target: Enhanced performance in Reason v4 evaluations, focusing on generating responses that elevate teacher-side Reason scores.
- Training Method: Offline DPO on Reason-ranked duel pairs, distinct from SFT or online GRPO.
- Context Length: Utilizes a
max_lenof 6144 tokens, indicating a focus on shorter contexts. - LoRA Configuration: Trained with LoRA r=32 (MidRank) and α=128 (HiAlpha).
- Performance: Achieved a margin of +0.005461 against the live king reign36 under wvk=7, leading to a "WIN / Stage-5 licensed" decision.
Intended Use
This model is specifically designed for:
- SN120 Affine miner submissions.
- Evalsrv Reason v4 duels.
It is not intended as a general-purpose chat model.