saidutta69/Qwen3-0.6B-heretic

TEXT GENERATIONConcurrent Unit Cost:1Model Size:0.8BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jul 2, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

saidutta69/Qwen3-0.6B-heretic is a 0.6 billion parameter variant of the Qwen3 model, created by RACER IS OP using the Heretic v1.2.0 abliteration method. This model is specifically engineered to suppress refusal behavior through targeted weight edits, maintaining the base model's knowledge and instruction-following capabilities. It is optimized for use cases requiring a Qwen3 model without refusal guardrails, suitable for CPU-only inference, edge deployment, or studying refusal mechanisms.

Loading preview...

Qwen3-0.6B-heretic: Decensored Qwen3 Variant

This model, developed by RACER IS OP, is a 0.6 billion parameter version of the Qwen3 architecture, modified using the Heretic v1.2.0 abliteration technique. Its primary distinction is the deliberate suppression of refusal behavior through targeted weight edits, rather than fine-tuning. This approach ensures that the base Qwen3 model's core knowledge and instruction-following abilities, including its unique thinking/non-thinking dual-mode, remain largely intact.

Key Capabilities and Features

  • Refusal Suppression: Achieves a significant reduction in refusal responses (from 56/100 to 5/100) compared to the original Qwen/Qwen3-0.6B model.
  • Minimal Divergence: Exhibits an exceptionally low KL divergence of 0.0018, indicating that the modifications are highly targeted and preserve the base model's characteristics.
  • Dual-Mode Architecture: Retains the Qwen3's ability to operate in both <think> and direct-answer modes.
  • Lightweight: At 0.6B parameters, it is suitable for resource-constrained environments.

Ideal Use Cases

  • CPU-only Inference & Edge Deployment: Its small size makes it efficient for local and embedded applications.
  • Research on Refusal Mechanisms: Provides a valuable testbed for studying and understanding how refusal behaviors are implemented and suppressed in reasoning-capable models.
  • Applications Requiring Unfiltered Responses: Suitable for developers who need a Qwen3 model that will comply with requests the base model would typically refuse, with the understanding that no safety filtering is layered on top. Users are responsible for its deployment and outputs.