elevateecho/sn120-03affd41ac7c
The elevateecho/sn120-03affd41ac7c model is a 35.1 billion parameter Affine SN120 challenger for Reason v4, developed by elevateecho. This model was trained using offline DPO on Reason-ranked duel pairs, optimizing for thoughts that raise teacher-side Reason. With a context length of 8192 tokens, it is specifically intended for SN120 Affine miner submissions and evalsrv Reason v4 duels, rather than general chat applications.
Loading preview...
Model Overview
This model, elevateecho/sn120-03affd41ac7c, is a 35.1 billion parameter Affine SN120 challenger specifically designed for Reason v4 (weight_version_key=7). It employs a tempered multi-sample log-mean-exp over k=3 teacher references.
Training Details
The model was trained using offline DPO (Direct Preference Optimization) on Reason-ranked duel pairs, rather than SFT or online GRPO. The optimization focused on preferring thoughts that enhance teacher-side Reason. It was fine-tuned from the vera6/affine-5g4yy75zuz-t6 base model using 604 filtered duel preference pairs. Key hyperparameters include a LoRA rank of 32, alpha of 128, beta of 0.1, and a learning rate of 2e-6. It supports a maximum context length of 8192 tokens.
Performance & Validation
During its development, the model demonstrated a margin of +0.006632 against the live king reign36 under wvk=7, with a z-score of 2.042 over 79 samples. It achieved a B pass rate of 0.521, meeting the required threshold. This checkpoint was validated as a WIN / Stage-5 licensed model.
Intended Use
This model is not a general chat model. Its primary and intended use is for SN120 Affine miner submissions and evalsrv Reason v4 duels.