unconstai/1787207581
The unconstai/1787207581 model is a 35.1 billion parameter Affine SN120 challenger, fine-tuned using offline DPO on Reason-ranked duel pairs. It is specifically optimized for improving 'Reason v4' scores in evaluation server duels, demonstrating a significant margin over its base model. This model is designed for specialized mining submissions and evaluations rather than general chat applications, featuring a 32768 token context length.
Loading preview...
Model Overview
unconstai/1787207581, also known as R959, is a 35.1 billion parameter Affine SN120 challenger model. It was developed by unconstai and is specifically designed to improve performance on 'Reason v4' evaluations. This model is not a general-purpose chat model but rather a specialized tool for competitive mining and evaluation server duels.
Training Methodology
R959 was trained using offline DPO (Direct Preference Optimization) on Reason-ranked duel pairs, rather than Supervised Fine-Tuning (SFT) or online GRPO. The optimization focused on increasing preference for thoughts that enhance the teacher-side Reason score. Key training parameters included a LoRA rank of 64 (HiRank), an alpha of 128 (HiAlpha), a beta of 0.1 (MidBeta), and a very low learning rate of 5e-7 (UltraLoLR). It was trained for 4 epochs over 19200 steps, utilizing a substantial context length of 12288 tokens (SoftCtx).
Performance and Differentiation
Compared to its base model, vera6/affine-5g4yy75zuz-t6, R959 achieved a significant margin of +0.006384 on Reason v4 evaluations, with a z-score of 2.451. This indicates a statistically significant improvement in its intended evaluation metric. The model's lineage traces back through several specialized R-series models, emphasizing its unique focus on specific evaluation criteria rather than broad language understanding.
Intended Use
This model is explicitly intended for SN120 Affine miner submissions and evalsrv Reason v4 duels. It is not recommended for use as a general chat model due to its highly specialized fine-tuning objective.