saidutta69/Qwen3.5-2B-heretic

VISIONConcurrent Unit Cost:1Model Size:2.3BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 2, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

saidutta69/Qwen3.5-2B-heretic is a 2.3 billion parameter decensored variant of the Qwen3.5-2B model, developed by saidutta69 using the Heretic v1.4.0 abliteration method. This model suppresses refusal behavior via targeted weight edits, leaving the base model's knowledge and instruction-following capabilities largely intact. It is optimized for developers seeking a fast Qwen3.5 model without guardrails, capable of running on CPU and edge devices with strong speed and long-context handling for its size.

Loading preview...

Overview

saidutta69/Qwen3.5-2B-heretic is a 2.3 billion parameter language model derived from Qwen/Qwen3.5-2B. Its primary distinction is the deliberate suppression of refusal behaviors, achieved through a technique called abliteration (directional ablation) using the Heretic v1.4.0 tool. Unlike traditional fine-tuning, abliteration directly edits specific weight directions responsible for refusals, preserving the base model's core knowledge and instruction-following abilities.

Key Capabilities

  • Decensored Output: Designed to comply with requests that the original Qwen3.5-2B model would typically refuse, without additional safety filtering.
  • Efficient Performance: The 2B hybrid linear-attention core is optimized for fast execution on CPU and edge devices.
  • Long Context Handling: Maintains strong performance with long context lengths, suitable for various applications.
  • Intact Base Model Capabilities: Retains the knowledge and instruction-following prowess of the original Qwen3.5-2B, as abliteration avoids degrading coherence often seen with fine-tuning for helpfulness.

Good For

  • Developers requiring a Qwen3.5 model without built-in refusal guardrails.
  • Applications where the base model's knowledge and instruction-following are critical, but safety filters are managed externally.
  • Deployment on resource-constrained environments like CPUs and edge devices due to its efficient architecture and size.

This model is provided with a clear warning regarding responsible use, as its refusal suppression is intentional and means it will generate content the base model would have declined.