reaperdoesntknow/Qwen3-0.6B-Distilled-30B-A3B-Thinking-SFT

Hugging Face
TEXT GENERATIONConcurrent Unit Cost:1Model Size:0.8BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Mar 22, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Warm

reaperdoesntknow/Qwen3-0.6B-Distilled-30B-A3B-Thinking-SFT is a 0.6 billion parameter Qwen3 causal language model developed by Convergent Intelligence LLC. It was created through a two-stage process: knowledge distillation from a 30B-parameter 'Thinking' teacher model for reasoning, followed by supervised fine-tuning on legal instruction data. This model is optimized for ultra-lightweight reasoning and instruction-following in legal and STEM domains, designed to run efficiently on edge devices with a small footprint.

Loading preview...

Model Overview

This model, developed by Convergent Intelligence LLC, is a 0.6 billion parameter Qwen3-based causal language model. It is distinguished by its unique two-stage training approach: initial knowledge distillation from a 30B-parameter 'Thinking' teacher model to establish a structured reasoning backbone, followed by supervised fine-tuning on legal instruction data. This methodology aims to transfer deep reasoning structures efficiently, achieving a 50x compression ratio from its teacher model.

Key Capabilities

  • Structured Reasoning: Leverages a 'Thinking' teacher model's extended deliberation traces to instill a robust reasoning structure, particularly in STEM domains.
  • Legal Instruction-Following: Fine-tuned on legal datasets, enabling it to perform legal analysis by applying learned derivation structures.
  • Ultra-Lightweight Deployment: At 0.6B parameters and under 500MB when quantized, it is designed for efficient inference on mobile, edge, and IoT devices.
  • Proof-Weighted Distillation: Utilizes a novel loss function that prioritizes reasoning steps over answer formatting during distillation.

Good For

  • Ultra-lightweight reasoning tasks on resource-constrained devices.
  • Instruction-following in legal and STEM contexts.
  • Educational tutoring applications requiring structured derivation.
  • Integration as a component in multi-model pipelines where small size is critical.

Limitations

Due to its compact size, the model has inherent capacity constraints. It may exhibit reasoning errors that larger models would not, and multi-step derivations beyond approximately 8 steps can degrade performance. While capable of general legal concepts, it lacks the nuance of larger models, and performance is weakest on underrepresented domains like molecular biology and physiology. Outputs should always be verified.