saidutta69/Qwen3.5-9B-heretic

VISIONPricing:Input $0.431 / Cached $0.0862 / Output $1.12Concurrent Unit Cost:1Model Size:9BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jul 21, 2026License:otherArchitecture:Transformer Featherless Exclusive Cold

saidutta69/Qwen3.5-9B-heretic is a 9 billion parameter large language model based on the Qwen3.5 architecture, featuring a 32768 token context length. Developed by saidutta69 using Heretic v1.4.0, this model is a decensored variant of Qwen/Qwen3.5-9B, with refusal behaviors suppressed via targeted weight edits rather than fine-tuning. It maintains the base model's strong multilingual reasoning and agentic capabilities, making it suitable for developers requiring a powerful LLM without built-in refusal guardrails.

Loading preview...

Overview

saidutta69/Qwen3.5-9B-heretic is a 9 billion parameter large language model derived from Qwen/Qwen3.5-9B. Its primary distinction is the suppression of refusal behaviors through "abliteration" using Heretic v1.4.0. This method involves targeted weight edits to the attention output and MLP down-projections, preserving the base model's knowledge and instruction-following capabilities while removing its inherent refusal guardrails. The model supports a 32768 token context length.

Key Capabilities & Features

  • Decensored Output: Designed to comply with requests that the base Qwen3.5-9B model would typically refuse, without degrading coherence.
  • Retained Base Model Utility: The abliteration process ensures that the base model's strong multilingual reasoning and agentic capabilities remain largely intact.
  • Hardware Accessibility: Provided with a full suite of GGUF quantizations (Q4_K_M, Q5_K_M, Q6_K, Q8_0, F16), enabling efficient deployment on consumer-grade GPUs (e.g., RTX 3060/4070 with 12 GB VRAM) and even CPU-only systems.
  • Broad Compatibility: Loads natively in llama.cpp, Ollama, LM Studio, Jan, vLLM, and SGLang.

Performance & Differentiation

Evaluations show a refusal rate of 79/100 on adversarial prompts, compared to 86/100 for the base Qwen3.5-9B. A KL divergence of 0.0015 from the base model indicates that the refusal suppression is extremely narrow, with minimal impact on the overall output distribution and utility. This model is not a capability upgrade but offers the same performance as Qwen3.5-9B without its refusal mechanisms.

Responsible Use

Users are advised that this model will comply with requests the base model would refuse, including potentially harmful ones, as there is no safety filtering layered on top. Users are responsible for its deployment and usage.