YYYYYYibo/qwen3-4b-questa-nonmcq-estep-pos20-neg20-truncneg20-step50

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 22, 2026Architecture:Transformer Featherless Exclusive Cold

YYYYYYibo/qwen3-4b-questa-nonmcq-estep-pos20-neg20-truncneg20-step50 is a 4 billion parameter Qwen3-based research model fine-tuned for mathematical problem-solving. This experimental checkpoint, derived from Qwen3-4B-Instruct-2507, focuses on non-multiple-choice math problems using a teacher-penalty PPO signal. It is specifically optimized for generating correct solutions and minimizing truncation in mathematical reasoning tasks, demonstrating improved accuracy in in-sample diagnostics.

Loading preview...

Overview

This model, YYYYYYibo/qwen3-4b-questa-nonmcq-estep-pos20-neg20-truncneg20-step50, is a 4 billion parameter research checkpoint based on the Qwen3 architecture. It was initialized from Qwen/Qwen3-4B-Instruct-2507 and represents an experimental E-step in a research project.

Key Capabilities & Training

  • Mathematical Problem Solving: Specifically fine-tuned on 5,897 deduplicated non-multiple-choice math problems from the YYYYYYibo/QuestA-OpenR1-Math-220k-Joined dataset.
  • Reinforcement Learning: Utilizes a sampled teacher-penalty PPO signal, corresponding to KL(teacher || frozen student), combined with sequence-level correctness rewards.
  • Reward Structure: Employs a +20 reward for correct answers, a -20 penalty for incorrect answers, and a -20 penalty for length truncation.
  • Context Handling: Supports a maximum prompt length of 4,096 tokens and a maximum training response length of 8,192 tokens.

Performance Insights (In-Sample Diagnostics)

Diagnostic evaluations on 256 randomly selected training problems show:

  • Base Model (no solution): 56.25% Math-Verify accuracy.
  • Base Model (with solution): 92.19% Math-Verify accuracy.
  • This Checkpoint (with solution): Achieved 78.91% Math-Verify accuracy, with a 2.73% truncation rate. This indicates a significant improvement over the base model without a solution, though it's an experimental artifact and not a generalization result.

Intended Use

This model is primarily a research artifact for exploring reinforcement learning techniques in mathematical reasoning. It is suitable for researchers and developers interested in fine-tuning LLMs for complex, non-multiple-choice mathematical problem generation and verification, particularly those focusing on reducing truncation and improving solution accuracy.