unconst/Affine-5czsc2fc98-r637-r252-odpo-midrank-lobeta-softctx-ep3-lolr-merged

TEXT GENERATIONPricing:Input $0.4 / Cached $0.07 / Output $4Concurrent Unit Cost:3Model Size:35.1BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Aug 17, 2026Architecture:Transformer Featherless Exclusive Cold

The unconst/Affine-5czsc2fc98-r637-r252-odpo-midrank-lobeta-softctx-ep3-lolr-merged model is a 35.1 billion parameter Affine SN120 challenger, derived from the unconst/Affine-5czsc2fc98-r252-merged base. It was trained using offline DPO on Reason-ranked pairs, specifically optimizing for higher teacher-side Reason scores. This model utilizes a soft context band (max_len=12288), mid LoRA rank (r=32), and low DPO beta (0.02) with a low learning rate (1e-6) over 3 epochs. Its primary intended use is for SN120 Affine miner submissions and evalsrv Reason duels, rather than general chat applications.

Loading preview...

Model Overview

This model, unconst/Affine-5czsc2fc98-r637-r252-odpo-midrank-lobeta-softctx-ep3-lolr-merged, is a 35.1 billion parameter Affine SN120 challenger. It is built upon the unconst/Affine-5czsc2fc98-r252-merged base and was specifically trained to excel in "Reason v3" tasks, a teacher-anchored scoring system.

Training Methodology

The model was developed using offline DPO (Direct Preference Optimization), not Supervised Fine-Tuning (SFT) or online GRPO. The optimization focused on improving preference for higher teacher-side Reason scores on mined pairs. Key training parameters include:

  • Data: SoftCtx × MidRank × LoBeta pair set, indicating a soft context band, mid LoRA rank, and low DPO beta.
  • LoRA Configuration: r=32 (MidRank) and α=128 (HiAlpha).
  • DPO Beta: β=0.02 (LoBeta).
  • Learning Rate: lr=1e-6 (LoLR).
  • Context Length: max_len=12288 (SoftCtx).
  • Epochs: Trained for 3 epochs over 3600 steps.

Performance & Differentiation

This model demonstrated a positive margin of +0.005735 against the live king reign34 (cryptoDev23/Affine-5Dku3dYp9j-hk8161), with a z-score of 2.91 over 77 samples. It achieved a thought median of 168.5 and a B pass rate of 0.4125. This specific training regimen and focus on Reason v3 tasks differentiate it from general-purpose language models.

Intended Use

This model is explicitly designed for SN120 Affine miner submissions and evalsrv Reason duels. It is not intended as a general chat model.