reaperdoesntknow/Qwen3-0.6B-Distilled-30B-A3B-Thinking-SFT

Hugging Face
TEXT GENERATIONPricing:Input $0.04 / Cached $0.008 / Output $0.08Concurrent Unit Cost:1Model Size:0.8BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Mar 22, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Warm

reaperdoesntknow/Qwen3-0.6B-Distilled-30B-A3B-Thinking-SFT is a 0.6 billion parameter Qwen3-based causal language model developed by Convergent Intelligence LLC. It was created through a two-stage process: knowledge distillation from a 30B-parameter "Thinking" teacher model for structured reasoning, followed by supervised fine-tuning on legal instruction data. This model is optimized for ultra-lightweight reasoning in legal and STEM domains, designed to run efficiently on edge devices with a small footprint (under 500MB quantized). Its primary strength lies in transferring deep reasoning structures for instruction-following tasks, particularly in legal analysis and STEM derivations.

Loading preview...

Model Overview

This model, reaperdoesntknow/Qwen3-0.6B-Distilled-30B-A3B-Thinking-SFT, is a compact 0.6 billion parameter Qwen3-based language model developed by Convergent Intelligence LLC. It employs a unique two-stage training approach to achieve advanced reasoning capabilities within a highly compressed footprint, approximately 50x smaller than its teacher model. The model's core innovation lies in its ability to distill complex reasoning structures from a 30B-parameter "Thinking" teacher model, which generates extended internal reasoning traces, before specializing in legal instruction-following.

Key Capabilities

  • Structured Reasoning: Built on a STEM chain-of-thought reasoning backbone, enabling rigorous derivation and logical chaining.
  • Legal Instruction-Following: Fine-tuned on legal datasets to apply its reasoning structure to legal analysis.
  • Extreme Compression: Achieves 50x compression from its teacher, allowing it to run on resource-constrained devices like mobile phones (under 500MB quantized).
  • Proof-Weighted Distillation: Utilizes a novel loss function that prioritizes reasoning steps over answer formatting during distillation, enhancing structural understanding.

Training Methodology

The model's training involved two distinct stages:

  1. Knowledge Distillation: From a Qwen3-30B-A3B-Thinking teacher model using 6,122 STEM chain-of-thought samples. This stage focused on transferring the teacher's deliberation and reasoning paths.
  2. Supervised Fine-Tuning: On the Alignment-Lab-AI/Lawyer-Instruct dataset, leveraging the established reasoning backbone for legal applications.

Good for

  • Ultra-lightweight reasoning on mobile, edge, and IoT devices.
  • Legal and STEM instruction-following tasks.
  • Educational tutoring and embedded inference applications.
  • Use as a component in multi-model pipelines where a small, reasoning-capable model is needed.