pb09204048/Qwen3-8B-DAPO-iter449-disable-thinking

TEXT GENERATIONPricing:Input $0.468 / Output $1.82Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Aug 12, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

This is an 8.2 billion parameter Qwen3 model, fine-tuned by pb09204048 using GRPO/DAPO reinforcement learning on the DAPO-Math-17k dataset. Optimized specifically for mathematical reasoning in non-thinking mode, it demonstrates significant performance gains on AIME-24 and AIME-25 benchmarks compared to the base Qwen3-8B model. It is designed for efficient, specialized mathematical problem-solving without explicit reasoning traces.

Loading preview...

Overview

This model, pb09204048/Qwen3-8B-DAPO-iter449-disable-thinking, is an 8.2 billion parameter Qwen3 variant. It has been fine-tuned using GRPO/DAPO reinforcement learning on the DAPO-Math-17k dataset, specifically with Qwen3's chat template in non-thinking mode (enable_thinking=false). This means the model was trained without exposure to <think>...</think> reasoning traces.

Key Capabilities & Performance

  • Enhanced Mathematical Reasoning: Achieves substantial improvements on AIME-24 and AIME-25 benchmarks in non-thinking mode, with gains of +48.75 pp and +49.18 pp respectively over the base Qwen3-8B model.
  • Efficient Problem Solving: Designed to provide direct answers without generating explicit reasoning steps, making it suitable for applications where only the final solution is required.
  • Qwen3 Foundation: Inherits core capabilities from the Qwen3 series, including strong instruction-following and multilingual support, though this specific checkpoint is optimized for a specialized mathematical task.

When to Use This Model

  • Mathematical Problem Solving: Ideal for tasks requiring accurate mathematical answers where the intermediate reasoning process is not needed or should be suppressed.
  • Efficiency-Focused Applications: Suitable for scenarios demanding quick, direct responses without the overhead of generating thinking traces.
  • Non-Thinking Mode Preference: If your application specifically benefits from a model that operates without explicit internal thought processes, this fine-tuned version is highly relevant.