thealper2/qwen3-0.6b-reasoning-sft
Thealper2/qwen3-0.6b-reasoning-sft is a 0.6 billion parameter Qwen3-based language model fine-tuned for enhanced reasoning capabilities. Developed by thealper2, this model specializes in generating step-by-step thought processes within a native block, followed by a final answer. It is optimized for tasks requiring methodical problem-solving and mathematical verification, leveraging a 32K context length.
Loading preview...
Model Overview
This model, thealper2/qwen3-0.6b-reasoning-sft, is a supervised fine-tune of the Qwen/Qwen3-0.6B base model, specifically designed to improve its reasoning abilities. It has 0.6 billion parameters and was trained using a full fine-tuning approach with TRL's SFTTrainer.
Key Capabilities
- Enhanced Reasoning: Fine-tuned on the
Kenshiii/minimax-m3-reasoning-tracesdataset, which emphasizes step-by-step problem-solving. - Structured Reasoning Output: Generates thought processes within a Qwen3 native
<think>…</think>block, followed by the final answer, making its reasoning transparent. - Mathematical Problem Solving: Demonstrated capability in solving and verifying mathematical equations, as shown in the usage example.
- Qwen3 Chat Template: Utilizes the standard Qwen3 chat template, ensuring compatibility with existing Qwen3 workflows.
Training Details
The model was trained for approximately 1.51 epochs (80 steps) on a dataset formatted to map reasoning traces to Qwen3's reasoning_content. It achieved a best validation loss of 1.0698. The training process used fp32 master weights and bfloat16 autocast, with a training sequence length of 2048 tokens.
Good for
- Applications requiring transparent, step-by-step reasoning.
- Tasks involving mathematical problem-solving and verification.
- Developers looking for a compact model with improved logical deduction over its base counterpart.