pb09204048/Qwen3-8B-DAPO-iter449-disable-thinking
This is an 8.2 billion parameter Qwen3 model, fine-tuned by pb09204048 using GRPO/DAPO reinforcement learning on the DAPO-Math-17k dataset. Optimized specifically for mathematical reasoning in non-thinking mode, it demonstrates significant performance gains on AIME-24 and AIME-25 benchmarks compared to the base Qwen3-8B model. It is designed for efficient, specialized mathematical problem-solving without explicit reasoning traces.
Loading preview...
Overview
This model, pb09204048/Qwen3-8B-DAPO-iter449-disable-thinking, is an 8.2 billion parameter Qwen3 variant. It has been fine-tuned using GRPO/DAPO reinforcement learning on the DAPO-Math-17k dataset, specifically with Qwen3's chat template in non-thinking mode (enable_thinking=false). This means the model was trained without exposure to <think>...</think> reasoning traces.
Key Capabilities & Performance
- Enhanced Mathematical Reasoning: Achieves substantial improvements on AIME-24 and AIME-25 benchmarks in non-thinking mode, with gains of +48.75 pp and +49.18 pp respectively over the base Qwen3-8B model.
- Efficient Problem Solving: Designed to provide direct answers without generating explicit reasoning steps, making it suitable for applications where only the final solution is required.
- Qwen3 Foundation: Inherits core capabilities from the Qwen3 series, including strong instruction-following and multilingual support, though this specific checkpoint is optimized for a specialized mathematical task.
When to Use This Model
- Mathematical Problem Solving: Ideal for tasks requiring accurate mathematical answers where the intermediate reasoning process is not needed or should be suppressed.
- Efficiency-Focused Applications: Suitable for scenarios demanding quick, direct responses without the overhead of generating thinking traces.
- Non-Thinking Mode Preference: If your application specifically benefits from a model that operates without explicit internal thought processes, this fine-tuned version is highly relevant.