crazyape777/mir-unconst-affine-5czsc2fc98-r1032-vera-odp
The crazyape777/mir-unconst-affine-5czsc2fc98-r1032-vera-odp is a 35.1 billion parameter Affine SN120 challenger model, derived from the vera6/affine-5g4yy75zuz-t6 base. It was trained using offline DPO on Reason-ranked duel pairs, specifically optimized for preference towards thoughts that enhance teacher-side Reason. With a short context length of 6144 tokens, this model is intended for SN120 Affine miner submissions and evalsrv Reason v4 duels, rather than general chat applications.
Loading preview...
Model Overview
The crazyape777/mir-unconst-affine-5czsc2fc98-r1032-vera-odp model, also known as R1032, is a 35.1 billion parameter Affine SN120 challenger. It is specifically designed for Reason v4 evaluations, utilizing a tempered multi-sample log-mean-exp over k=3 teacher references. This model is a specialized checkpoint, not a general-purpose chat model.
Training Methodology
This model was trained using offline DPO (Direct Preference Optimization) on Reason-ranked duel pairs, rather than Supervised Fine-Tuning (SFT) or online GRPO. The optimization objective was to enhance preference for thoughts that increase teacher-side Reason scores. Key training parameters include:
- Base Model:
vera6/affine-5g4yy75zuz-t6 - LoRA Configuration: r=32 (MidRank), α=128 (HiAlpha)
- DPO Beta: β=0.3 (HiBeta)
- Learning Rate: lr=1e-6 (MidLR)
- Context Length: max_len=6144 (ShortCtx)
- Data: 604 lines of ShortCtx filtered duel preference pairs (
dpo_duel_reason.jsonl)
Performance and Intended Use
During local evaluation against the live king vera6/affine-5g4yy75zuz-t6 under wvk=7, R1032 demonstrated a positive margin of +0.005461, indicating a win. Its primary intended use is for SN120 Affine miner submissions and evalsrv Reason v4 duels. It is explicitly stated not to be a general chat model.