saidutta69/Qwen3.5-2B-heretic

VISIONPricing:Input $0.32 / Cached $0.064 / Output $1.6Concurrent Unit Cost:1Model Size:2.3BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 2, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The saidutta69/Qwen3.5-2B-heretic is a 2.3 billion parameter decensored variant of the Qwen/Qwen3.5-2B model, developed by saidutta69 using the Heretic v1.4.0 abliteration method. This model suppresses refusal behavior through targeted weight edits, preserving the base model's knowledge and instruction-following capabilities. It is optimized for developers seeking a fast Qwen3.5 model without refusals, offering strong speed and long-context handling on CPU and edge devices.

Loading preview...

Overview

The saidutta69/Qwen3.5-2B-heretic is a 2.3 billion parameter language model derived from the Qwen/Qwen3.5-2B base model. Its primary distinction lies in its decensored nature, achieved through the "abliteration" method using Heretic v1.4.0. This technique involves targeted weight edits to the attention output and MLP down-projections to suppress refusal behavior, rather than fine-tuning. This approach ensures that the base model's core knowledge and instruction-following abilities remain largely intact, providing a model that complies with requests the original might refuse.

Key Capabilities

  • Decensored Output: Deliberately suppresses refusal behavior, allowing the model to respond to prompts that the base Qwen3.5-2B would typically decline.
  • Preserved Base Capabilities: Maintains the original Qwen3.5-2B's knowledge and instruction-following, as abliteration avoids the coherence degradation often seen with fine-tuning for helpfulness.
  • Efficient Performance: The 2B hybrid linear-attention core is designed for speed and long-context handling, suitable for deployment on CPU and edge devices.
  • Extensive GGUF Support: Provides a full ladder of GGUF quantizations, enabling flexible deployment across various hardware configurations, from high-end GPUs to CPU-only setups.

Good For

  • Developers requiring a fast Qwen3.5 model without built-in refusal guardrails.
  • Applications where direct, unfiltered responses are preferred or necessary.
  • Deployment on resource-constrained environments like CPUs or edge devices, thanks to its efficient architecture and GGUF quantizations.