PursuitOfDataScience/Qwen3-0.6b-thinking

TEXT GENERATIONPricing:Input $0.04 / Cached $0.008 / Output $0.08Concurrent Unit Cost:1Model Size:0.8BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Feb 17, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

PursuitOfDataScience/Qwen3-0.6b-thinking is a 0.8 billion parameter language model, a Chain-of-Thought (CoT) Supervised Fine-Tuned (SFT) version of Qwen/Qwen3-0.6B. It is specifically trained to perform step-by-step reasoning using explicit / traces before generating a final answer. This model excels at mathematical and reasoning tasks, demonstrating a significant improvement on the GSM8K benchmark compared to its base model. It is optimized for applications requiring transparent, verifiable reasoning processes.

Loading preview...

PursuitOfDataScience/Qwen3-0.6b-thinking: Chain-of-Thought Reasoning

This model is a 0.8 billion parameter Chain-of-Thought (CoT) Supervised Fine-Tuned (SFT) version of the Qwen/Qwen3-0.6B base model. Its core innovation lies in its training to explicitly generate step-by-step reasoning within <think>...</think> tags before providing a final answer, making its thought process transparent.

Key Capabilities & Features

  • Explicit Step-by-Step Reasoning: Generates detailed intermediate thoughts, enhancing interpretability and verifiability of answers.
  • Improved Mathematical Reasoning: Achieves a +13.88 percentage-point improvement on the GSM8K Pass@1 benchmark (from 28.73% to 42.61%) compared to its base model.
  • Specialized Training Data: Fine-tuned on the PursuitOfDataScience/MiniMax-M2.1-Mixture-of-Thoughts dataset, which contains rich reasoning traces.
  • Optimized for Inference: Uses PyTorch SDPA and torch.compile for efficient generation.

Use Cases & Considerations

This model is particularly well-suited for tasks requiring logical deduction, problem-solving, and mathematical reasoning where understanding the 'how' behind the answer is crucial. Its explicit reasoning traces can be valuable for debugging, auditing, or educational applications. While it shows strong performance for its size, its 0.6B parameter count means its reasoning depth is inherently limited compared to much larger models. Further training could potentially yield even greater accuracy.