reaperdoesntknow/Qwen3-1.7B-Thinking-Distil
reaperdoesntknow/Qwen3-1.7B-Thinking-Distil is a 2.03 billion parameter Qwen3ForCausalLM model developed by Convergent Intelligence LLC. It is distilled from the Qwen3-30B-A3B-Thinking teacher model, specifically capturing extended deliberation and reasoning patterns. With a 40,960 token context length, this model excels at generating long-form reasoning chains and internal monologues, making it highly effective for complex analytical tasks.
Loading preview...
Overview
reaperdoesntknow/Qwen3-1.7B-Thinking-Distil is a 2.03 billion parameter language model from Convergent Intelligence LLC, distilled from the larger Qwen3-30B-A3B-Thinking teacher. This model uniquely captures and compresses the teacher's extended deliberation patterns, including reasoning through uncertainty, backtracking, and re-evaluation, into a compact 1.7B parameter student model. It was fine-tuned using Supervised Fine-Tuning (SFT) on the longwriter-6k dataset, which provides long-form generation samples rich in reasoning chains.
Key Capabilities
- Extended Reasoning: Specializes in generating detailed, multi-step reasoning processes and internal monologues, reflecting how a larger model might "think" through a problem.
- Deliberative Depth: Designed to produce rich signal in its output, showcasing reconsideration and resolution before arriving at a conclusion.
- Efficient Size: Achieves advanced reasoning capabilities within a 1.7B effective parameter count, making it efficient for deployment.
- Long Context: Supports a maximum position context length of 40,960 tokens, enabling generation of extensive reasoning outputs.
Good For
- Applications requiring detailed, step-by-step explanations or problem-solving narratives.
- Tasks where understanding the "thought process" behind an answer is as important as the answer itself.
- Generating content that mimics human-like deliberation, re-evaluation, and complex analytical thinking.
- Use cases benefiting from a smaller model that retains the sophisticated reasoning of a much larger teacher model.