best26/Affine-megaextra
best26/Affine-megaextra is a 35.1 billion parameter Affine SN120 challenger model, developed by best26, specifically trained using offline DPO to optimize for reasoning tasks. It was fine-tuned on Reason-ranked pairs with a SoftCtx and HiRank dataset, aiming to surpass previous models in reasoning performance. This model is intended for specialized mining submissions and evaluation server reasoning duels, rather than general chat applications.
Loading preview...
Affine-megaextra: Optimized for Reasoning
best26/Affine-megaextra is a 35.1 billion parameter model developed by best26, designed as an SN120 challenger to excel in reasoning tasks. It was trained using an offline DPO (Direct Preference Optimization) method, specifically targeting the optimization of preference for higher teacher-side Reason scores on mined pairs.
Key Training Details
- Base Model:
unconst/Affine-5czsc2fc98-r252-merged. - Optimization Target: Preference for superior reasoning performance, as measured by
lpC(y_C|z_A) − lpC(y_C|∅). - Data: Utilized a SoftCtx × HiRank pair set with a MidBeta (β=0.1) configuration, indicating a focus on specific context lengths and high-rank data filtering.
- Hyperparameters: Notable settings include LoRA r=64, α=128, lr=5e-6, and a maximum context length of 12288 tokens.
- Performance: Achieved a margin of +0.006196 against the live king reign34 model in local n80 evaluations, demonstrating its improved reasoning capabilities.
Intended Use
This model is not a general chat model. Its primary purpose is for Affine miner submissions and evaluation server reasoning duels, where its specialized optimization for reasoning can be leveraged effectively.