YYYYYYibo/qwen3-4b-questa-nonmcq-estep-pos20-neg20-truncneg20-step50
YYYYYYibo/qwen3-4b-questa-nonmcq-estep-pos20-neg20-truncneg20-step50 is a 4 billion parameter Qwen3-based research model fine-tuned for mathematical problem-solving. This experimental checkpoint, derived from Qwen3-4B-Instruct-2507, focuses on non-multiple-choice math problems using a teacher-penalty PPO signal. It is specifically optimized for generating correct solutions and minimizing truncation in mathematical reasoning tasks, demonstrating improved accuracy in in-sample diagnostics.
Loading preview...
Overview
This model, YYYYYYibo/qwen3-4b-questa-nonmcq-estep-pos20-neg20-truncneg20-step50, is a 4 billion parameter research checkpoint based on the Qwen3 architecture. It was initialized from Qwen/Qwen3-4B-Instruct-2507 and represents an experimental E-step in a research project.
Key Capabilities & Training
- Mathematical Problem Solving: Specifically fine-tuned on 5,897 deduplicated non-multiple-choice math problems from the
YYYYYYibo/QuestA-OpenR1-Math-220k-Joineddataset. - Reinforcement Learning: Utilizes a sampled teacher-penalty PPO signal, corresponding to
KL(teacher || frozen student), combined with sequence-level correctness rewards. - Reward Structure: Employs a
+20reward for correct answers, a-20penalty for incorrect answers, and a-20penalty for length truncation. - Context Handling: Supports a maximum prompt length of 4,096 tokens and a maximum training response length of 8,192 tokens.
Performance Insights (In-Sample Diagnostics)
Diagnostic evaluations on 256 randomly selected training problems show:
- Base Model (no solution): 56.25% Math-Verify accuracy.
- Base Model (with solution): 92.19% Math-Verify accuracy.
- This Checkpoint (with solution): Achieved 78.91% Math-Verify accuracy, with a 2.73% truncation rate. This indicates a significant improvement over the base model without a solution, though it's an experimental artifact and not a generalization result.
Intended Use
This model is primarily a research artifact for exploring reinforcement learning techniques in mathematical reasoning. It is suitable for researchers and developers interested in fine-tuning LLMs for complex, non-multiple-choice mathematical problem generation and verification, particularly those focusing on reducing truncation and improving solution accuracy.