elevateecho/sn120-f20753584e79
The elevateecho/sn120-f20753584e79 model is a 35.1 billion parameter Affine SN120 challenger, fine-tuned using offline DPO on Reason-ranked duel pairs. It is specifically optimized to raise teacher-side Reason scores, utilizing a tempered multi-sample log-mean-exp over three teacher references. This model is intended for use as an SN120 Affine miner submission and for evalsrv Reason v4 duels, rather than as a general chat model. It features a SoftCtx context length of 12288 tokens and was trained with LoRA r=32 and alpha=128.
Loading preview...
Model Overview
This model, elevateecho/sn120-f20753584e79, is an Affine SN120 challenger, specifically a Reason v4 variant. It is a 35.1 billion parameter model derived from the vera6/affine-5g4yy75zuz-t6 base.
Key Characteristics
- Training Method: Offline DPO (Direct Preference Optimization) on Reason-ranked duel pairs, not SFT or online GRPO.
- Optimization Goal: Optimized to enhance preferences for thoughts that increase teacher-side Reason scores.
- Context Length: Utilizes a SoftCtx maximum length of 12288 tokens.
- LoRA Configuration: Trained with LoRA r=32 (MidRank) and α=128 (HiAlpha).
- Performance: Achieved a margin of +0.003665 with a z-score of 2.177 over the live king reign36 under
wvk=7, leading to a WIN / Stage-5 licensed decision.
Intended Use
This model is designed as an SN120 Affine miner submission and for evalsrv Reason v4 duels. It is explicitly not intended as a general chat model, focusing instead on its specialized optimization for Reason-based evaluations.