best26/Affine-5czsc2fc98-r861-vera-odpo-midrank-midbeta-softctx-megaextra-ep4-ultralolr-merged

TEXT GENERATIONConcurrent Unit Cost:3Model Size:35.1BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Aug 19, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

best26/Affine-5czsc2fc98-r861-vera-odpo-midrank-midbeta-softctx-megaextra-ep4-ultralolr-merged is a 35.1 billion parameter Affine model fine-tuned using offline DPO on Reason-ranked duel pairs. Developed by best26, this model is specifically optimized for the SN120 Affine miner submission and evalsrv Reason v4 duels. It focuses on generating thoughts that enhance teacher-side Reason scores, distinguishing it from general chat models. With a context length of 32768 tokens, it is designed for specialized competitive AI environments rather than broad conversational applications.

Loading preview...

Model Overview

best26/Affine-5czsc2fc98-r861-vera-odpo-midrank-midbeta-softctx-megaextra-ep4-ultralolr-merged is a 35.1 billion parameter Affine model, specifically an SN120 challenger for Reason v4. It was developed by best26 and is a specialized checkpoint derived from vera6/affine-5g4yy75zuz-t6.

Training Methodology

This model was trained using offline DPO (Direct Preference Optimization) on Reason-ranked duel pairs, rather than Supervised Fine-Tuning (SFT) or online GRPO. The optimization objective was to enhance preferences for thoughts that increase the teacher-side Reason score, utilizing a tempered multi-sample log-mean-exp over k=3 teacher references. The training data consisted of Soft Mid Mid Soft \u00d7 SoftCtx filtered duel preference pairs.

Key Training Parameters & Features

  • LoRA Configuration: r=32 (MidRank), \u03b1=128 (HiAlpha)
  • DPO Beta: \u03b2=0.1 (MidBeta)
  • Learning Rate: 5e-7 (UltraLoLR)
  • Context Length: 12288 tokens (SoftCtx) during training, with the model supporting 32768 tokens.
  • Training Steps: 19200 steps over 4 epochs.
  • Performance: Achieved a margin of +0.003665 with a z-score of 2.177 against the live king reign36 under wvk=7, leading to a WIN / Stage-5 licensed decision.

Intended Use

This model is explicitly designed for SN120 Affine miner submissions and evalsrv Reason v4 duels. It is not intended as a general chat model and its performance is optimized for competitive AI reasoning tasks within its specific domain.