crazyape777/mir-829-odpo-midrank-lobeta-ep3

TEXT GENERATIONPricing:Input $0.4 / Cached $0.07 / Output $4Concurrent Unit Cost:3Model Size:35.1BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Aug 17, 2026Architecture:Transformer Featherless Exclusive Cold

crazyape777/mir-829-odpo-midrank-lobeta-ep3 is a 35.1 billion parameter Affine model, developed by crazyape777, specifically trained using offline DPO to optimize for reasoning tasks. This model is a specialized SN120 challenger designed to excel in Reason v3 evaluations, making it suitable for competitive AI mining and evaluation services. It was trained with a low DPO beta and a low learning rate on a SoftCtx × MidRank × LoBeta pair set, focusing on preference for higher teacher-side Reason scores.

Loading preview...

Model Overview

crazyape777/mir-829-odpo-midrank-lobeta-ep3 is a 35.1 billion parameter Affine model, developed by crazyape777, specifically engineered as an SN120 challenger. Its primary objective is to outperform existing models on Reason v3 evaluations, which measure a model's reasoning capabilities based on a teacher-anchored scoring system.

Training Methodology

This checkpoint was trained using offline DPO (Direct Preference Optimization), distinguishing it from traditional SFT (Supervised Fine-Tuning) or online GRPO methods. The optimization focused on enhancing the model's preference for higher teacher-side Reason scores derived from mined pairs. Key training details include:

  • Base Model: unconst/Affine-5czsc2fc98-r252-merged
  • Data: A specialized "SoftCtx × MidRank × LoBeta" pair set, indicating soft context banding, mid LoRA rank, and low DPO beta.
  • Hyperparameters: Notable settings include LoRA r=32 (MidRank), α=128 (HiAlpha), a low β=0.02 (LoBeta), and a low learning rate (lr=1e-6, LoLR). The maximum context length was 12288 tokens (SoftCtx), trained over 3 epochs with 3600 steps.

Performance & Intended Use

During internal evaluations, this model demonstrated a positive margin of +0.005735 against the live king reign34 (cryptoDev23/Affine-5Dku3dYp9j-hk8161), with a z-score of 2.91 over 77 samples. Its thought median was 168.5 (≥80) and B pass rate was 0.4125 (≥0.30).

Intended Use: This model is specifically designed for SN120 Affine miner submissions and evalsrv Reason duels. It is not intended as a general chat model but rather as a specialized tool for competitive AI reasoning tasks and evaluations.