Phase-Technologies/qwen2.5-3b-claude-distilled-reasoning-dpo
TEXT GENERATIONConcurrent Unit Cost:1Model Size:3.1BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jul 31, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold
Phase-Technologies/qwen2.5-3b-claude-distilled-reasoning-dpo is a 3.09 billion parameter causal language model based on the Qwen2.5 architecture. It is post-trained using supervised fine-tuning on Claude 3.5 Sonnet reasoning traces and Direct Preference Optimization. This model specializes in step-by-step mathematical, logical, and code-synthesis reasoning, offering factually grounded responses with a low memory footprint.
Loading preview...
Overview
This model, qwen2.5-3b-claude-distilled-reasoning-dpo, is a 3.09 billion parameter causal language model built upon Qwen/Qwen2.5-3B-Instruct. It undergoes a two-stage post-training alignment process to enhance its reasoning capabilities and factual accuracy.
Key Capabilities
- Distilled Chain-of-Thought (CoT): Excels at step-by-step reasoning for mathematical equations, coding challenges, and logic puzzles, derived from Claude 3.5 Sonnet monologue traces.
- Factually Grounded: Demonstrates high performance on graduate/research-level physics and mathematics questions, significantly reducing hallucinations common in basic SFT models.
- Alignment: Utilizes Direct Preference Optimization (DPO) with
argilla/ultrafeedback-binarized-preferences-cleanedto suppress scientific hallucinations and token repetition loops. - Low Memory Footprint: Designed to run efficiently in FP16/SDPA, requiring approximately 6GB VRAM on a single consumer GPU.
- ChatML Ready: Fully compatible with standard Qwen2.5 ChatML chat templates and system prompt instructions.
Good For
- Applications requiring robust, step-by-step reasoning in mathematics, logic, and code synthesis.
- Tasks where factual accuracy in scientific and technical domains is critical.
- Deployment on resource-constrained hardware, such as single consumer GPUs, due to its optimized memory usage.