ledgernova/sn120-3b427b1678d5
The ledgernova/sn120-3b427b1678d5 model is a 35.1 billion parameter Affine SN120 challenger developed by ledgernova, specifically optimized for Reason v4 duels. It was trained using offline DPO on Reason-ranked duel pairs, focusing on preferences that raise teacher-side Reason. This model is intended for use as an SN120 Affine miner submission and evalsrv Reason v4 duel, rather than a general chat model. It features a SoftCtx context length of 12288 tokens and utilizes HiRank LoRA with r=64 and alpha=128.
Loading preview...
Overview
ledgernova/sn120-3b427b1678d5 is a 35.1 billion parameter Affine SN120 model developed by ledgernova, designed as a challenger for Reason v4 (weight_version_key=7). It employs a tempered multi-sample log-mean-exp over k=3 teacher references. The model's core mechanism for evaluating turns is defined by a_i = lpC(y_i|z_A) − lpC(y_i|∅), with Reason = τ·log(mean_i exp(a_i/τ)). It also requires a median stripped thought length of |z|≥80 and a B pass of ≥0.30.
Training Methodology
This checkpoint was trained using offline DPO (Direct Preference Optimization) on Reason-ranked duel pairs, rather than SFT or online GRPO. The optimization focused on preferences for thoughts that enhance teacher-side Reason. Key hyperparameters include:
- LoRA: r=64 (HiRank), α=128 (HiAlpha)
- Beta: β=0.1 (MidBeta)
- Learning Rate: lr=5e-7 (UltraLoLR)
- Context Length: max_len=12288 (SoftCtx)
- Steps/Epochs: max_steps=19200, epochs=4
Performance and Intended Use
The model demonstrated a significant margin of +0.006384 (SE 0.002605, z=2.451, n=79) against the live king vera6/affine-5g4yy75zuz-t6@8e3f1695e058837ed80fec3238ff439fdc2d0f0e under wvk=7, leading to a WIN / Stage-5 licensed decision. Its thought median was 173 and B pass 0.406. This model is specifically intended for SN120 Affine miner submissions and evalsrv Reason v4 duels, and is not a general chat model.