elevateecho/sn120-3a778f1fd066
The elevateecho/sn120-3a778f1fd066 model is a 35.1 billion parameter Affine SN120 challenger, trained using offline DPO on Reason-ranked pairs. It is specifically optimized to excel in reasoning tasks, demonstrating a significant margin over its predecessor, the r252 model. This model is primarily intended for SN120 Affine miner submissions and evaluation server Reason duels, rather than general chat applications.
Loading preview...
Model Overview
The elevateecho/sn120-3a778f1fd066 model, also known as R596, is an Affine SN120 challenger. It was developed to surpass the performance of the live king model on Reason v3 benchmarks, utilizing a teacher-anchored scoring method. This model is a direct successor to the unconst/Affine-5czsc2fc98-r252-merged base model.
Training Methodology
This checkpoint was trained using offline DPO (Direct Preference Optimization), specifically on Reason-ranked pairs. Unlike SFT (Supervised Fine-Tuning) or online GRPO, the optimization focused on enhancing preference for higher teacher-side Reason scores on mined pairs. The training data comprised a SoftCtx × HiRank pair set with a MidBeta β=0.1, and a soft context length band with a high-rank filter. Key hyperparameters included LoRA r=64, α=128, and a learning rate of 5e-6, with a maximum sequence length of 12288 tokens.
Performance Highlights
In local n80 evaluations against the live king reign34 (cryptoDev23/Affine-5Dku3dYp9j-hk8161), the model achieved a margin of +0.006196 with a standard error of 0.002357, indicating a strong performance in reasoning tasks. It also demonstrated a thought median of 199 (≥80) and a B pass rate of 0.368 (≥0.30).
Intended Use
This model is specifically designed for SN120 Affine miner submissions and evaluation server Reason duels. It is not intended as a general chat model.