JaeWooShin/qwen2.5-math-1.5b-fewgt-u1-n256-inverted-lam1-s0
JaeWooShin/qwen2.5-math-1.5b-fewgt-u1-n256-inverted-lam1-s0 is a 1.5 billion parameter Qwen2.5-Math model, anchored to original base weights, and trained for 300 GRPO steps with a 32768 token context length. This research artifact was specifically trained with an inverted reward function that pays for wrong answers, making it a probe for understanding model behavior under spurious rewards. Its training was confined to a rank-one band, penalizing weight changes outside a single direction, and it is not intended for general-purpose deployment.
Loading preview...
Model Overview
This model, JaeWooShin/qwen2.5-math-1.5b-fewgt-u1-n256-inverted-lam1-s0, is a 1.5 billion parameter Qwen2.5-Math variant. It is explicitly a research artifact and not a general-purpose model, as it was trained with a unique, inverted reward function that incentivizes incorrect answers in mathematical problems. The primary goal of this model is to investigate how a language model behaves when rewarded for being wrong, serving as a mechanism probe rather than a functional tool.
Key Training Characteristics
- Inverted Reward: The model was trained to receive a reward of 1.0 for incorrect parseable
\boxed{}answers and 0.0 for correct ones, consuming up to 19,200 ground-truth labels for reward computation. - Rank-One Band Constraint: The update process was confined to a rank-one band, meaning weight changes were penalized if they lay outside a single, pre-determined direction (
u1). Thisu1direction was derived from a separate donor run with true correctness rewards. - Anchored Training: Training started and remained anchored to the original base weights (W0) of
Qwen/Qwen2.5-Math-1.5B.
Observed Performance
Despite being rewarded for wrong answers, the model showed an increase in "accuracy" (meaning, it became better at producing wrong answers that were rewarded) on various math benchmarks like AIME, AMC, MATH500, Minerva, and OlympiadBench, compared to the base W0 model. For instance, on MATH500, its accuracy (of being wrong) increased from 0.296 at step 0 to 0.602 at step 300.
Limitations
- Not for General Use: This model is not designed for deployment or general use due to its training objective of rewarding incorrectness.
- Untested Behavior: Its behavior outside of short-form boxed math is untested, and no checks for safety, instruction following, or general capabilities were performed.
- Research-Specific: It represents a single experimental setup (single seed, model, label count, lambda) and lacks control arms for direct comparison of the band's effect versus other setup components.