bluecolor777/megaextra-ep4-ultralolr-merged
The bluecolor777/megaextra-ep4-ultralolr-merged model is a 35.1 billion parameter Affine SN120 challenger, specifically optimized for Reason v4 evaluations. Trained using offline DPO on Reason-ranked duel pairs, it focuses on generating thoughts that enhance teacher-side Reason scores. With a context length of 32768 tokens, this model is intended for specialized mining submissions and evaluation server duels, rather than general chat applications.
Loading preview...
Model Overview
The bluecolor777/megaextra-ep4-ultralolr-merged is a 35.1 billion parameter Affine SN120 challenger model, specifically developed for Reason v4 evaluations. It is not a general-purpose chat model but is highly specialized for competitive mining submissions and evaluation server duels.
Key Training Details
- Base Model:
vera6/affine-5g4yy75zuz-t6(live king reign36). - Methodology: Trained using offline DPO (Direct Preference Optimization) on Reason-ranked duel pairs, rather than SFT or online GRPO. This approach optimizes for preferences that raise teacher-side Reason scores.
- Data: Utilized Soft Mid Mid Soft × SoftCtx filtered duel preference pairs from
dpo_duel_reason.jsonl. - Hyperparameters: Notable settings include LoRA r=32 (MidRank), α=128 (HiAlpha), β=0.1 (MidBeta), and an ultra-low learning rate (lr=5e-7, UltraLoLR).
- Context Length: Supports a maximum context length of 12288 tokens during training (SoftCtx).
Performance & Validation
The model demonstrated a positive margin of +0.003665 against the live king reign36 under wvk=7 evaluations, achieving a z-score of 2.177 over 80 samples. It met the required thought median (141.5, ≥80) and B pass rate (0.5375, ≥0.30), leading to a WIN / Stage-5 licensed decision.
Intended Use
This model is specifically designed for SN120 Affine miner submissions and evalsrv Reason v4 duels. It is not intended for use as a general chat model.