thealper2/qwen3-0.6b-reasoning-sft

TEXT GENERATIONPricing:Input $0.04 / Cached $0.008 / Output $0.08Concurrent Unit Cost:1Model Size:0.8BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 24, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

Thealper2/qwen3-0.6b-reasoning-sft is a 0.6 billion parameter Qwen3-based language model fine-tuned for enhanced reasoning capabilities. Developed by thealper2, this model specializes in generating step-by-step thought processes within a native block, followed by a final answer. It is optimized for tasks requiring methodical problem-solving and mathematical verification, leveraging a 32K context length.

Loading preview...

Model Overview

This model, thealper2/qwen3-0.6b-reasoning-sft, is a supervised fine-tune of the Qwen/Qwen3-0.6B base model, specifically designed to improve its reasoning abilities. It has 0.6 billion parameters and was trained using a full fine-tuning approach with TRL's SFTTrainer.

Key Capabilities

  • Enhanced Reasoning: Fine-tuned on the Kenshiii/minimax-m3-reasoning-traces dataset, which emphasizes step-by-step problem-solving.
  • Structured Reasoning Output: Generates thought processes within a Qwen3 native <think>…</think> block, followed by the final answer, making its reasoning transparent.
  • Mathematical Problem Solving: Demonstrated capability in solving and verifying mathematical equations, as shown in the usage example.
  • Qwen3 Chat Template: Utilizes the standard Qwen3 chat template, ensuring compatibility with existing Qwen3 workflows.

Training Details

The model was trained for approximately 1.51 epochs (80 steps) on a dataset formatted to map reasoning traces to Qwen3's reasoning_content. It achieved a best validation loss of 1.0698. The training process used fp32 master weights and bfloat16 autocast, with a training sequence length of 2048 tokens.

Good for

  • Applications requiring transparent, step-by-step reasoning.
  • Tasks involving mathematical problem-solving and verification.
  • Developers looking for a compact model with improved logical deduction over its base counterpart.