pb09204048/Qwen3-8B-DAPO-iter599-thinking

TEXT GENERATIONPricing:Input $0.468 / Output $1.82Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Aug 21, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

This is an 8.2 billion parameter Qwen3 model, fine-tuned by pb09204048 using GRPO/DAPO reinforcement learning on the DAPO-Math-17K dataset. It is specifically trained in 'thinking mode' to generate internal thought processes before providing a final answer, distinguishing it from models that directly output responses. The model supports a 32,768 token context length natively, extendable to 131,072 tokens with YaRN, and is optimized for complex logical reasoning, mathematics, and agentic tasks.

Loading preview...

Model Overview

This model, pb09204048/Qwen3-8B-DAPO-iter599-thinking, is an 8.2 billion parameter Qwen3 variant fine-tuned using GRPO/DAPO reinforcement learning on the DAPO-Math-17K dataset. Its key differentiator is training in 'thinking mode' (enable_thinking=true), meaning it generates an internal <think>...</think> block before the final answer, designed to enhance complex logical reasoning.

Key Capabilities & Differentiators

  • Thinking Mode: Uniquely trained to produce explicit thought processes, which can be beneficial for debugging and understanding the model's reasoning steps in complex tasks like mathematics and coding.
  • Reinforcement Learning: Fine-tuned with GRPO/DAPO using the slime framework, focusing on improving performance on mathematical problems.
  • Context Length: Natively supports a 32,768 token context length, extendable up to 131,072 tokens using the YaRN method for processing long texts.
  • Agentic Use: Excels in tool-calling capabilities, recommended for use with Qwen-Agent for integrating external tools.
  • Multilingual Support: Part of the Qwen3 series, which supports over 100 languages and dialects.

Should You Use This Model?

This model is particularly suited for use cases requiring transparent reasoning and complex problem-solving, especially in mathematics and logical tasks, where the intermediate thought process is valuable. If your application benefits from understanding how the model arrived at an answer, or if you are working on agentic systems that leverage internal reasoning, this 'thinking mode' variant offers a distinct advantage over models that only provide direct outputs. For general-purpose dialogue or scenarios where efficiency without explicit reasoning steps is paramount, other Qwen3 variants (e.g., those trained with disable-thinking) might be more appropriate.