tastegarden/sn120-62dfb322fdce

TEXT GENERATIONPricing:Input $0.4 / Cached $0.07 / Output $4Concurrent Unit Cost:3Model Size:35.1BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Aug 20, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The tastegarden/sn120-62dfb322fdce is a 35.1 billion parameter Affine SN120 challenger model, fine-tuned using offline DPO on Reason-ranked duel pairs. It is specifically optimized for the Reason v4 evaluation server duel, demonstrating a positive margin against the live king reign36. This model is designed for specialized evaluation tasks rather than general chat applications, focusing on improving teacher-side Reason scores.

Loading preview...

Model Overview

tastegarden/sn120-62dfb322fdce is a 35.1 billion parameter Affine SN120 model, developed as a challenger for the Reason v4 evaluation server. It was trained using an offline DPO (Direct Preference Optimization) method, specifically on Reason-ranked duel pairs, rather than traditional SFT or online GRPO. The model's optimization focused on enhancing preference for thoughts that elevate teacher-side Reason scores, utilizing a tempered multi-sample log-mean-exp over three teacher references.

Key Training Details

  • Base Model: vera6/affine-5g4yy75zuz-t6@8e3f1695e058837ed80fec3238ff439fdc2d0f0e (live king reign36).
  • Data: ShortCtx filtered duel preference pairs (dpo_duel_reason.jsonl, 604 lines).
  • Hyperparameters: Notable settings include LoRA r=32 (MidRank), α=128 (HiAlpha), β=0.3 (HiBeta), learning rate of 1e-6 (MidLR), and a maximum context length of 6144 tokens (ShortCtx).
  • Performance: Achieved a margin of +0.005461 with a z-score of 2.098 against the live king reign36 under wvk=7, indicating a WIN / Stage-5 licensed decision.

Intended Use

This model is specifically intended as an SN120 Affine miner submission and for evalsrv Reason v4 duel evaluations. It is not designed or recommended as a general chat model.