crazyape777/mir-unconst-affine-5czsc2fc98-r861-vera-odpo
The crazyape777/mir-unconst-affine-5czsc2fc98-r861-vera-odpo model is a 35.1 billion parameter Affine SN120 challenger, fine-tuned using offline DPO on Reason-ranked duel pairs. It is specifically optimized for improving 'Reason v4' performance by preferring thoughts that enhance teacher-side reasoning. This model is intended for specialized use in Affine mining submissions and evalsrv Reason v4 duels, rather than general chat applications.
Loading preview...
Model Overview
This model, crazyape777/mir-unconst-affine-5czsc2fc98-r861-vera-odpo, is an Affine SN120 challenger derived from the vera6/affine-5g4yy75zuz-t6 base model. It was trained using an offline DPO (Direct Preference Optimization) method, specifically optimizing for preferences that increase 'Reason v4' scores by committing to a teacher next-action mode.
Key Training Details
- Methodology: Offline DPO on Reason-ranked duel pairs, distinct from SFT or online GRPO.
- Optimization Target: Enhanced preference for thoughts that elevate teacher-side Reason scores, utilizing a tempered multi-sample log-mean-exp over k=3 teacher references.
- Data: Filtered duel preference pairs from
dpo_duel_reason.jsonlwith specificSoft Mid Mid Soft × SoftCtxfiltering. - Hyperparameters: Notable settings include LoRA r=32 (MidRank), α=128 (HiAlpha), β=0.1 (MidBeta), and a very low learning rate (lr=5e-7, UltraLoLR).
- Context Length: Trained with a maximum sequence length of 12288 tokens (SoftCtx).
- Performance: Achieved a positive margin of +0.003665 against the live king
reign36underwvk=7, leading to its Stage-5 licensing.
Intended Use
This model is not a general chat model. Its primary purpose is for:
- Affine SN120 miner submissions.
- Evalsrv Reason v4 duels.
It is highly specialized for its intended optimization target and specific evaluation framework.