affine-kraus/Affine-5czsc2fc98-r1032-vera-odpo-midrank-hibeta-shortctx-ultraextra-ep4-midlr-merged
Affine-5czsc2fc98-r1032-vera-odpo-midrank-hibeta-shortctx-ultraextra-ep4-midlr-merged is a 35.1 billion parameter model developed by Affine, derived from the vera6/affine-5g4yy75zuz-t6 base. This model was trained using offline DPO to optimize for 'Reason v4' performance in duel scenarios, specifically preferring thoughts that enhance teacher-side reasoning. With a context length of 32768 tokens, it is primarily intended for SN120 Affine miner submissions and evaluation server Reason v4 duels, rather than general chat applications.
Loading preview...
Model Overview
affine-kraus/Affine-5czsc2fc98-r1032-vera-odpo-midrank-hibeta-shortctx-ultraextra-ep4-midlr-merged is a 35.1 billion parameter model developed by Affine, specifically designed as an SN120 challenger for Reason v4. It is based on the vera6/affine-5g4yy75zuz-t6 model.
Training Methodology
This checkpoint was trained using offline DPO (Direct Preference Optimization) on Reason-ranked duel pairs, rather than Supervised Fine-Tuning (SFT) or online GRPO. The optimization objective was to enhance preference for thoughts that increase teacher-side Reason scores, utilizing a tempered multi-sample log-mean-exp over k=3 teacher references. Key training parameters include:
- LoRA: r=32 (MidRank), α=128 (HiAlpha)
- Beta: β=0.3 (HiBeta)
- Learning Rate: lr=1e-6 (MidLR)
- Max Context Length: 6144 tokens (ShortCtx) during training, though the model supports 32768 tokens.
- Epochs: 4
Performance and Intended Use
The model demonstrated a positive margin of +0.005461 against the live king vera6/affine-5g4yy75zuz-t6 under wvk=7, indicating a WIN / Stage-5 licensed status. Its primary intended use is for SN120 Affine miner submissions and evaluation server Reason v4 duels. It is not designed as a general chat model.