pandora-box/Affine-5eqdtdzqle-ko8dsnxa
Affine-5eqdtdzqle-ko8dsnxa is a 35.1 billion parameter model developed by pandora-box, fine-tuned using offline DPO on Reason-ranked duel pairs. This model is specifically optimized for improving teacher-side Reason scores in SN120 Affine miner submissions and evalsrv Reason v4 duels. It is not intended as a general-purpose chat model but rather for specialized evaluation and mining tasks.
Loading preview...
Model Overview
Affine-5eqdtdzqle-ko8dsnxa is a 35.1 billion parameter model from pandora-box, developed as an SN120 challenger for Reason v4 (weight_version_key=7). It utilizes a tempered multi-sample log-mean-exp over three teacher references. This checkpoint was trained via offline DPO (Direct Preference Optimization) on Reason-ranked duel pairs, specifically optimizing for preferences that enhance teacher-side Reason scores.
Key Training Details
- Base Model:
vera6/affine-5g4yy75zuz-t6 - Methodology: Offline DPO, not SFT (Supervised Fine-Tuning) or online GRPO.
- Optimization Target: Preference for thoughts that increase teacher-side Reason.
- Data: Soft Mid Mid Soft × SoftCtx filtered duel preference pairs from
dpo_duel_reason.jsonl. - Hyperparameters: Notable settings include LoRA r=32 (MidRank), α=128 (HiAlpha), β=0.1 (MidBeta), and a low learning rate (lr=5e-7, UltraLoLR). It was trained for 4 epochs with a
max_lenof 12288 (SoftCtx).
Performance & Intended Use
This model demonstrated a positive margin of +0.003665 against the live king vera6/affine-5g4yy75zuz-t6 under wvk=7, leading to a WIN / Stage-5 licensed decision. It is explicitly intended for SN120 Affine miner submissions and evalsrv Reason v4 duels, and is not designed as a general chat model.