pandora-box/Affine-5eqdtdzqle-ko8dsnxa

TEXT GENERATIONConcurrent Unit Cost:3Model Size:35.1BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Aug 19, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

Affine-5eqdtdzqle-ko8dsnxa is a 35.1 billion parameter model developed by pandora-box, fine-tuned using offline DPO on Reason-ranked duel pairs. This model is specifically optimized for improving teacher-side Reason scores in SN120 Affine miner submissions and evalsrv Reason v4 duels. It is not intended as a general-purpose chat model but rather for specialized evaluation and mining tasks.

Loading preview...

Model Overview

Affine-5eqdtdzqle-ko8dsnxa is a 35.1 billion parameter model from pandora-box, developed as an SN120 challenger for Reason v4 (weight_version_key=7). It utilizes a tempered multi-sample log-mean-exp over three teacher references. This checkpoint was trained via offline DPO (Direct Preference Optimization) on Reason-ranked duel pairs, specifically optimizing for preferences that enhance teacher-side Reason scores.

Key Training Details

  • Base Model: vera6/affine-5g4yy75zuz-t6
  • Methodology: Offline DPO, not SFT (Supervised Fine-Tuning) or online GRPO.
  • Optimization Target: Preference for thoughts that increase teacher-side Reason.
  • Data: Soft Mid Mid Soft × SoftCtx filtered duel preference pairs from dpo_duel_reason.jsonl.
  • Hyperparameters: Notable settings include LoRA r=32 (MidRank), α=128 (HiAlpha), β=0.1 (MidBeta), and a low learning rate (lr=5e-7, UltraLoLR). It was trained for 4 epochs with a max_len of 12288 (SoftCtx).

Performance & Intended Use

This model demonstrated a positive margin of +0.003665 against the live king vera6/affine-5g4yy75zuz-t6 under wvk=7, leading to a WIN / Stage-5 licensed decision. It is explicitly intended for SN120 Affine miner submissions and evalsrv Reason v4 duels, and is not designed as a general chat model.