crazyape777/fk-unconst-affine-5czsc2fc98-r938-vera-odpo-midrank-hibeta-soft

TEXT GENERATIONPricing:Input $0.4 / Cached $0.07 / Output $4Concurrent Unit Cost:3Model Size:35.1BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Aug 20, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The crazyape777/fk-unconst-affine-5czsc2fc98-r938-vera-odpo-midrank-hibeta-soft model is a 35.1 billion parameter Affine SN120 challenger, fine-tuned using offline DPO on Reason-ranked duel pairs. It is optimized for improving preference for thoughts that raise teacher-side Reason, rather than general chat. This model features a 32768 token context length and utilizes LoRA with r=32 and α=128.

Loading preview...

Model Overview

The crazyape777/fk-unconst-affine-5czsc2fc98-r938-vera-odpo-midrank-hibeta-soft model is a 35.1 billion parameter Affine SN120 challenger, specifically developed for Reason v4 (weight_version_key=7). It is based on the vera6/affine-5g4yy75zuz-t6 parent model and was trained using an offline DPO (Direct Preference Optimization) method, not Supervised Fine-Tuning (SFT) or online GRPO.

Key Characteristics

  • Optimization Target: The model was optimized to enhance preference for thoughts that increase teacher-side Reason, using a tempered multi-sample log-mean-exp over k=3 teacher references.
  • Training Data: Utilized Soft Mid Mid Soft × SoftCtx filtered duel preference pairs from dpo_duel_reason.jsonl.
  • Hyperparameters: Notable settings include LoRA with r=32 (MidRank) and α=128 (HiAlpha), β=0.3 (HiBeta), a very low learning rate (lr=5e-7, UltraLoLR), and a substantial maximum context length of 12288 tokens (SoftCtx).
  • Performance: Achieved a win against the live king reign36 with a margin of +0.004951 and a z-score of 2.399 on 80 samples, meeting the criteria for Stage-5 licensing.

Intended Use

This model is specifically designed as an SN120 Affine miner submission and for evalsrv Reason v4 duels. It is not intended as a general-purpose chat model.