tastegarden/sn120-9591dd387030
The tastegarden/sn120-9591dd387030 model is a 35.1 billion parameter Affine SN120 challenger, specifically optimized for the Reason v4 evaluation metric. It was trained using offline DPO to enhance preference for thoughts that improve teacher-side Reason scores. This model features a SoftCtx context length of 12288 tokens and is intended for use in SN120 Affine miner submissions and evalsrv Reason v4 duels, rather than as a general chat model.
Loading preview...
Overview
This model, tastegarden/sn120-9591dd387030, is a 35.1 billion parameter Affine SN120 challenger, specifically fine-tuned for the Reason v4 evaluation metric. It was developed by tastegarden using an offline DPO (Direct Preference Optimization) method, distinct from SFT or online GRPO approaches. The training focused on optimizing for preferences that elevate teacher-side Reason scores, where filler content is penalized under a log-mean-exp aggregation.
Key Training Details
- Base Model:
vera6/affine-5g4yy75zuz-t6(live king reign36). - Methodology: Offline DPO on Reason-ranked duel pairs, specifically targeting thoughts that improve teacher-side Reason.
- Data: Filtered duel preference pairs from
dpo_duel_reason.jsonlwith Soft Mid Mid Soft × SoftCtx filtering. - Hyperparameters: Notable settings include LoRA r=32 (MidRank), α=128 (HiAlpha), β=0.1 (MidBeta), learning rate 5e-7 (UltraLoLR), and a maximum context length of 12288 tokens (SoftCtx).
- Performance: Achieved a margin of +0.003665 against the live king reign36 under wvk=7, with a z-score of 2.177, leading to a "WIN / Stage-5 licensed" decision.
Intended Use
This model is designed for SN120 Affine miner submissions and evalsrv Reason v4 duels. It is explicitly stated not to be a general chat model, indicating its specialized nature for specific evaluation and mining tasks.