reaperdoesntknow/Qwen3-1.7B-Thinking-Distil

Hugging Face
TEXT GENERATIONPricing:Input $0.32 / Cached $0.064 / Output $1.6Concurrent Unit Cost:1Model Size:2BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Mar 27, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Warm

The reaperdoesntknow/Qwen3-1.7B-Thinking-Distil is a 2 billion parameter Qwen3-based causal language model developed by Convergent Intelligence LLC. It is specifically fine-tuned to capture and reproduce extended reasoning chains and deliberative thought processes from a larger Qwen3-30B-A3B-Thinking teacher model. This model excels at generating long-form reasoning and internal monologues, making it suitable for tasks requiring detailed thought processes and problem-solving explanations.

Loading preview...

Model Overview

reaperdoesntknow/Qwen3-1.7B-Thinking-Distil is a 2 billion parameter Qwen3-based model from Convergent Intelligence LLC, specifically designed to distill and reproduce extended reasoning patterns. It was created by applying Supervised Fine-Tuning (SFT) on the longwriter-6k dataset, using the Qwen3-30B-A3B-Thinking model as a teacher. This teacher variant is known for generating detailed, long-form reasoning chains, including internal monologues, backtracking, and re-evaluation before reaching a conclusion.

Key Capabilities

  • Extended Reasoning: Captures the deliberative depth and internal monologue of a much larger teacher model, enabling it to reason through uncertainty and complex problems.
  • Long-Form Generation: Optimized for generating detailed explanations and thought processes, with a context length of 40,960 tokens and training on sequences up to 4,096 tokens.
  • Efficient Size: Compresses advanced reasoning capabilities into a compact 1.7B effective parameter model, making it efficient for deployment.
  • Structured Deliberation: Focuses on the structure of reasoning rather than just final answers, transferring how the teacher approaches and resolves problems.

Good For

  • Problem Solving: Generating step-by-step solutions or explanations for complex problems.
  • Educational Content: Creating detailed tutorials or explanations that mimic human thought processes.
  • Content Generation: Producing long-form text that requires deep reasoning and logical progression.
  • Research & Development: Exploring models that prioritize the process of thought over just the outcome, especially in resource-constrained environments.

For more advanced distillation techniques, refer to the TopologicalQwen project, which builds upon the foundational methodologies used here.