ShellFace/20260820-063650
ShellFace/20260820-063650 is a 35.1 billion parameter Affine SN120 challenger model, developed by ShellFace, specifically optimized for Reason v4 evaluations. It was trained using offline DPO on Reason-ranked duel pairs, focusing on generating thoughts that enhance teacher-side reasoning. With a context length of 32768 tokens, this model is designed for specialized evaluation server duels rather than general chat applications.
Loading preview...
Overview
ShellFace/20260820-063650 is a 35.1 billion parameter model, part of the Affine SN120 challenger series, specifically engineered for Reason v4 evaluations. It was developed by ShellFace and trained using an offline DPO (Direct Preference Optimization) method, distinct from SFT or online GRPO. The training focused on optimizing for preferences in thoughts that elevate teacher-side Reason scores, utilizing a tempered multi-sample log-mean-exp over k=3 teacher references.
Key Training Details
- Base Model:
vera6/affine-5g4yy75zuz-t6@8e3f1695e058837ed80fec3238ff439fdc2d0f0e - Optimization: Preference for thoughts that increase teacher-side Reason.
- Data: Soft Mid Mid Soft × SoftCtx filtered duel preference pairs.
- Hyperparameters: Notable settings include LoRA r=64 (HiRank), α=128 (HiAlpha), β=0.1 (MidBeta), and a maximum context length of 12288 tokens (SoftCtx).
- Performance: Achieved a margin of +0.006384 with a z-score of 2.451 against the live king reign36 under
wvk=7, indicating a successful challenge and Stage-5 licensing.
Intended Use
This model is specifically designed as an SN120 Affine miner submission for evalsrv Reason v4 duels. It is not intended for general chat applications but rather for specialized evaluation server tasks.