crazyape777/mir-unconst-affine-5czsc2fc98-r1008-vera-odp

TEXT GENERATIONPricing:Input $0.4 / Cached $0.07 / Output $4Concurrent Unit Cost:3Model Size:35.1BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Aug 20, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

crazyape777/mir-unconst-affine-5czsc2fc98-r1008-vera-odp is a 35.1 billion parameter Affine model, derived from vera6/affine-5g4yy75zuz-t6, specifically optimized for Reason v4 evaluations. This model was trained using offline DPO with a LoRA rank of 64 and an alpha of 128, focusing on improving preference for thoughts that enhance teacher-side Reason. It features a context length of 8192 tokens and is intended for SN120 Affine miner submissions and evalsrv Reason v4 duels, rather than general chat applications.

Loading preview...

Model Overview

crazyape777/mir-unconst-affine-5czsc2fc98-r1008-vera-odp is a 35.1 billion parameter Affine model, developed by crazyape777, specifically engineered as an SN120 Affine challenger for Reason v4 evaluations. It is based on the vera6/affine-5g4yy75zuz-t6 model and was trained using an offline DPO (Direct Preference Optimization) method, distinct from SFT or online GRPO. The primary optimization goal was to enhance the model's preference for responses that elevate teacher-side Reason scores, utilizing a tempered multi-sample log-mean-exp over three teacher references.

Key Training Details

  • Base Model: vera6/affine-5g4yy75zuz-t6
  • Methodology: Offline DPO on Reason-ranked duel pairs.
  • Optimization Target: Preference for thoughts that increase teacher-side Reason.
  • Hyperparameters:
    • LoRA r=64 (HiRank), α=128 (HiAlpha)
    • β=0.1 (MidBeta)
    • Learning Rate (lr)=1e-6 (MidLR)
    • Maximum context length (max_len)=8192 tokens (MidCtx)
    • Trained for 19200 steps over 4 epochs.

Performance and Intended Use

This model demonstrated a significant margin of +0.005917 with a z-score of 2.650 against the live king reign36 under wvk=7 evaluations, indicating a WIN / Stage-5 licensed status. Its design is highly specialized:

  • Primary Use Case: SN120 Affine miner submissions and evalsrv Reason v4 duels.
  • Limitation: It is not intended as a general-purpose chat model.