0xbidkslj1/base-model
The 0xbidkslj1/base-model, named R861, is a 35.1 billion parameter model developed by 0xbidkslj1. It is an Affine SN120 challenger specifically optimized for the Reason v4 evaluation metric, utilizing tempered multi-sample log-mean-exp over teacher references. This model was trained using offline DPO on Reason-ranked duel pairs, focusing on preferences for thoughts that enhance teacher-side Reason. Its primary intended use is for SN120 Affine miner submissions and evalsrv Reason v4 duels, rather than general chat applications.
Loading preview...
R861: An Affine SN120 Challenger for Reason v4
0xbidkslj1/base-model, codenamed R861, is a 35.1 billion parameter model specifically engineered as an Affine SN120 challenger for the Reason v4 evaluation metric. It employs a unique tempered multi-sample log-mean-exp approach over three teacher references (τ=0.03) to determine its Reason score. This model is not a general-purpose chat model but is highly specialized for competitive AI mining scenarios.
Training Methodology
R861 was developed from the vera6/affine-5g4yy75zuz-t6 base model using offline DPO (Direct Preference Optimization). Unlike SFT or online GRPO, its training optimized for preferences that increase the teacher-side Reason score, penalizing filler content. The training data consisted of SoftCtx filtered duel preference pairs, with key hyperparameters including LoRA r=32 (MidRank), α=128 (HiAlpha), β=0.1 (MidBeta), and a very low learning rate (lr=5e-7, UltraLoLR). It was trained for 4 epochs over 19200 steps with a maximum context length of 12288 tokens.
Performance and Intended Use
During its development, R861 demonstrated a positive margin of +0.003665 against the live king reign36 under wvk=7, achieving a z-score of 2.177. It successfully met the criteria for thought median (141.5 ≥ 80) and B pass rate (0.5375 ≥ 0.30), leading to its WIN / Stage-5 licensed decision. This model is explicitly designed for SN120 Affine miner submissions and evalsrv Reason v4 duels, making it suitable for competitive AI environments focused on specific reasoning tasks rather than broad conversational capabilities.