saidutta69/Qwen3.5-9B-heretic

VISIONConcurrent Unit Cost:1Model Size:9BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jul 21, 2026License:qwen-research-licenseArchitecture:Transformer Featherless Exclusive Cold

saidutta69/Qwen3.5-9B-heretic is a 9 billion parameter, 32K context length variant of Qwen/Qwen3.5-9B, created by saidutta69 using Heretic v1.4.0. This model features suppressed refusal behavior through targeted weight edits, maintaining the base model's strong multilingual reasoning and agentic capabilities. It is designed for developers seeking a Qwen model without built-in refusal guardrails, suitable for deployment on 12 GB GPUs or consumer hardware via GGUF quantizations.

Loading preview...

Model Overview

saidutta69/Qwen3.5-9B-heretic is a 9 billion parameter language model derived from Qwen/Qwen3.5-9B. Its primary distinction is the removal of refusal behaviors, achieved through a process called "abliteration" using Heretic v1.4.0. This method involves targeted weight edits to the attention output and MLP down-projections, which suppresses refusals while preserving the base model's core knowledge and instruction-following abilities.

Key Characteristics

  • Decensored Variant: Specifically engineered to remove refusal behaviors present in the base Qwen3.5-9B model.
  • Abliteration Method: Unlike fine-tuning, abliteration directly modifies weights responsible for refusal, aiming to maintain the original model's coherence and capabilities without degradation.
  • Base Model Capabilities: Inherits Qwen's strong multilingual reasoning and agentic capabilities.
  • Context Length: Supports a context length of 32,768 tokens.
  • Hardware Compatibility: Optimized for deployment on 12 GB GPUs or consumer hardware using various GGUF quantizations (Q4_K_M, Q5_K_M, Q6_K, Q8_0).

Use Cases

This model is intended for developers who require a Qwen-based model without built-in safety filtering or refusal mechanisms. It is suitable for applications where the user is responsible for content moderation and desires direct compliance with requests that the base model might otherwise refuse. It can be integrated with llama.cpp, transformers, Ollama, LM Studio, Jan, vLLM, and SGLang.