crazyape777/mir-822-odpo-hirank-midbeta
The crazyape777/mir-822-odpo-hirank-midbeta is a 35.1 billion parameter Affine model, developed by crazyape777, specifically optimized for reasoning tasks. It was trained using offline DPO on Reason-ranked pairs with a SoftCtx and HiRank dataset, focusing on improving preference for higher teacher-side Reason scores. This model is designed as an SN120 Affine miner submission for evaluation server Reason duels, rather than a general-purpose chat model.
Loading preview...
Overview
The crazyape777/mir-822-odpo-hirank-midbeta is a 35.1 billion parameter Affine model, developed by crazyape777, specifically engineered to excel in reasoning tasks. It serves as an SN120 Affine challenger, aiming to surpass existing models on the Reason v3 benchmark, which evaluates lpC(y_C|z_A) − lpC(y_C|∅) scores.
Key Training Details
- Base Model:
unconst/Affine-5czsc2fc98-r252-merged - Training Method: Offline DPO (Direct Preference Optimization) was applied, distinct from SFT or online GRPO.
- Optimization Target: The model was optimized to show a strong preference for higher teacher-side Reason scores on mined pairs.
- Data: Training utilized a
SoftCtx × HiRankpair set, incorporating a soft context length band and a high-rank filter, with a MidBeta (β=0.1) configuration. - Hyperparameters: Notable settings include LoRA r=64, α=128, learning rate 5e-6, and a maximum context length of 12288 tokens.
Performance Highlights
- Reason v3 Benchmark: Achieved a margin of +0.006196 against the live king reign34 model, with a z-score of 2.63 over 75 samples.
- Prior Comparison: Showed an even stronger margin of +0.008490 against its r252 predecessor.
Intended Use
- This model is primarily intended for SN120 Affine miner submissions and evaluation server Reason duels.
- It is not designed as a general-purpose chat model.