jacob-rojic/mirror-michael-chan-000-affine-5gbzvz7tcc-d1

TEXT GENERATIONPricing:Input $0.4 / Cached $0.07 / Output $4Concurrent Unit Cost:3Model Size:35.1BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Aug 22, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The jacob-rojic/mirror-michael-chan-000-affine-5gbzvz7tcc-d1 model is a 35.1 billion parameter Affine SN120 challenger, fine-tuned using offline DPO on Reason-ranked duel pairs. It is specifically optimized for the Reason v4 evaluation metric, employing a tempered multi-sample log-mean-exp over teacher references. This model is designed as a specialized miner submission for evaluation servers, rather than a general-purpose chat model, excelling in tasks that require advanced reasoning capabilities.

Loading preview...

Model Overview

The jacob-rojic/mirror-michael-chan-000-affine-5gbzvz7tcc-d1 model, internally designated R861, is a 35.1 billion parameter Affine SN120 challenger. It was developed by jacob-rojic as a specialized submission for the Reason v4 evaluation metric (weight_version_key=7). This model is not a general chat model but is highly optimized for specific reasoning tasks.

Training Methodology

This checkpoint was trained using offline DPO (Direct Preference Optimization) on Reason-ranked duel pairs, rather than Supervised Fine-Tuning (SFT) or online GRPO. The optimization focused on increasing preference for thoughts that elevate the teacher-side Reason score, utilizing a tempered multi-sample log-mean-exp over three teacher references. Key hyperparameters included LoRA with r=32 and α=128, a β of 0.1, and a learning rate of 5e-7, with a maximum context length of 12288 tokens.

Performance and Intended Use

The model demonstrated a positive margin of +0.003665 against the live king vera6/affine-5g4yy75zuz-t6 under wvk=7, with a B pass rate of 0.5375. It is primarily intended as an Affine miner submission for evaluation servers and for evalsrv Reason v4 duels. Users should note that its design is highly specialized for these benchmarks and not for general conversational AI applications.