ShellFace/20260817-213018
ShellFace/20260817-213018 is a 35.1 billion parameter Affine SN120 challenger model, specifically optimized for improving reasoning capabilities in a multi-sample log-mean-exp teacher reference system. Trained using offline DPO on Reason-ranked duel pairs, it focuses on enhancing thoughts that raise teacher-side Reason scores. This model is intended for specialized evaluation server duels and not as a general-purpose chat model.
Loading preview...
Overview
ShellFace/20260817-213018 is a 35.1 billion parameter model, part of the Affine SN120 challenger series, specifically designed to compete in Reason v4 duels. It utilizes a tempered multi-sample log-mean-exp over k=3 teacher references to evaluate and improve reasoning. The model was trained via offline DPO (Direct Preference Optimization) on preference pairs filtered for ShortCtx, MidRank, and HiBeta characteristics, focusing on optimizing for thoughts that enhance teacher-side Reason scores.
Key Training Details
- Base Model:
unconst/Affine-5czsc2fc98-r252-merged - Methodology: Offline DPO, not SFT or online GRPO.
- Optimization Target: Preference for thoughts that increase teacher-side Reason scores.
- Data: ShortCtx × MidRank × HiBeta filtered duel preference pairs.
- Hyperparameters: LoRA r=32, α=128, β=0.3, lr=1e-6, max_len=6144, max_steps=7200, epochs=3.
- Performance: Achieved a win in local n80 vs. live king reign34 evaluations with a margin of +0.002137 and a B pass rate of 0.304.
Intended Use
This model is specifically developed as an SN120 Affine miner submission for evaluation server Reason v4 duels. It is not intended for use as a general chat model.