reaperdoesntknow/Qwen3-0.6B-Distilled-30B-A3B-Thinking-SFT

Hugging Face
TEXT GENERATIONPricing:Input $0.04 / Cached $0.008 / Output $0.08Concurrent Unit Cost:1Model Size:0.8BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Mar 22, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Warm

reaperdoesntknow/Qwen3-0.6B-Distilled-30B-A3B-Thinking-SFT is a 0.6 billion parameter Qwen3-based causal language model developed by Convergent Intelligence LLC. It was created through a two-stage distillation process, first from a 30B-parameter 'Thinking' teacher for reasoning, then fine-tuned on legal instruction data. This model is optimized for ultra-lightweight reasoning and legal/STEM instruction-following, achieving a 50x compression ratio for deployment on edge devices.

Loading preview...

Model Overview

This model, reaperdoesntknow/Qwen3-0.6B-Distilled-30B-A3B-Thinking-SFT, is a compact 0.6 billion parameter Qwen3-based language model developed by Convergent Intelligence LLC. It is designed for efficient reasoning and specialized instruction-following, achieved through a unique two-stage training pipeline. The model first undergoes knowledge distillation from a 30B-parameter Qwen3-A3B-Thinking teacher, which generates extended internal reasoning traces, to establish a robust reasoning backbone. This is followed by supervised fine-tuning on legal instruction data, leveraging the structural similarities between mathematical and legal reasoning.

Key Capabilities

  • Efficient Reasoning: Distilled from a 'Thinking' teacher, it learns to generate internal reasoning traces, transferring deeper reasoning structures despite its small size.
  • Domain Specialization: Fine-tuned for both STEM chain-of-thought problems (Physics, Linear Algebra, Engineering, etc.) and legal instruction-following.
  • Ultra-Lightweight: At 0.6B parameters and under 500MB when quantized, it is suitable for mobile, edge, and IoT deployments.
  • Proof-Weighted Distillation: Utilizes a novel loss function that prioritizes reasoning steps over answer formatting during distillation, allocating limited capacity to structural understanding.

Good for

  • Ultra-lightweight reasoning on mobile, edge, or IoT devices.
  • Legal and STEM instruction-following tasks.
  • Educational tutoring applications.
  • Embedded inference and as a component in multi-model pipelines.
  • Scenarios requiring reasoning capabilities within strict memory constraints (under 500MB).