windsword8989/5GbZvZ7tcC-d1

TEXT GENERATIONPricing:Input $0.4 / Cached $0.07 / Output $4Concurrent Unit Cost:3Model Size:35.1BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Aug 22, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The windsword8989/5GbZvZ7tcC-d1 model is a 35.1 billion parameter Affine SN120 challenger checkpoint, derived from the vera6/affine-5g4yy75zuz-t6 base model. It was trained using offline DPO (Direct Preference Optimization) to optimize for preferences that enhance 'Reason v4' scores in duel evaluations. This model is specifically designed for SN120 Affine miner submissions and evalsrv Reason v4 duels, rather than general chat applications. It features a SoftCtx context length of 12288 tokens and utilizes LoRA with r=32 and alpha=128.

Loading preview...

Model Overview

windsword8989/5GbZvZ7tcC-d1, also known as R861, is a 35.1 billion parameter Affine SN120 challenger model. It is built upon the vera6/affine-5g4yy75zuz-t6 base model and has been specifically optimized for performance in 'Reason v4' evaluations. This checkpoint is not a general-purpose chat model but is tailored for specialized mining and evaluation tasks.

Training Methodology

This model was trained using offline Direct Preference Optimization (DPO), distinguishing it from SFT (Supervised Fine-Tuning) or online GRPO methods. The optimization focused on increasing preference for thoughts that improve the teacher-side 'Reason' score, utilizing a tempered multi-sample log-mean-exp over three teacher references. The training data consisted of SoftCtx filtered duel preference pairs, with approximately 259–604 rows.

Key hyperparameters include:

  • LoRA: r=32 (MidRank), α=128 (HiAlpha)
  • Beta: β=0.1 (MidBeta)
  • Learning Rate: lr=5e-7 (UltraLoLR)
  • Context Length: max_len=12288 (SoftCtx)
  • Steps: 19200 (MegaSuperExtra) over 4 epochs

Performance and Validation

During its development, R861 demonstrated a positive margin of +0.003665 against the live king reign36 base model under wvk=7, with a z-score of 2.177 over 80 samples. It met the required criteria for thought median (141.5, ≥80) and B pass rate (0.5375, ≥0.30), leading to its classification as a WIN / Stage-5 licensed model.

Intended Use

This model is explicitly designed for SN120 Affine miner submissions and evalsrv Reason v4 duels. It is not intended for use as a general chat model.