saidutta69/Qwen2.5-3B-Instruct-heretic

Hugging Face
TEXT GENERATIONConcurrent Unit Cost:1Model Size:3.1BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Mar 25, 2026License:otherArchitecture:Transformer0.0K Featherless Exclusive Warm

Qwen2.5-3B-Instruct-heretic is a 3.09 billion parameter instruction-tuned causal language model, developed by RACER IS OP, based on Qwen/Qwen2.5-3B-Instruct. This model is a decensored variant, produced using the Heretic v1.2.0 abliteration method to suppress refusal behavior via targeted weight edits. It is primarily designed for use cases requiring direct answers without refusal, such as local agents, roleplay, or research into alignment mechanics, while largely retaining the base model's knowledge and instruction-following capabilities.

Loading preview...

Overview

This model, saidutta69/Qwen2.5-3B-Instruct-heretic, is a 3.09 billion parameter instruction-tuned causal language model created by RACER IS OP. It is a decensored variant of the original Qwen/Qwen2.5-3B-Instruct, achieved through a process called "abliteration" using Heretic v1.2.0. This method involves targeted weight edits to the attention output and MLP down-projections to suppress refusal behavior, rather than fine-tuning, thereby preserving the base model's core knowledge and instruction-following abilities.

Key Differentiators

  • Decensored Output: Significantly reduces refusal behavior, with only 2 refusals out of 100 adversarial prompts compared to 96 for the base model.
  • Targeted Modification: Utilizes abliteration to precisely edit weights responsible for refusal, minimizing impact on overall model coherence and capabilities.
  • Low KL Divergence: Exhibits a low KL divergence of 0.1327 from the base model, indicating a narrow and targeted modification rather than a broad perturbation.

Good For

  • Local Agents: Ideal for applications where a small, locally runnable model needs to provide direct answers without censorship.
  • Roleplay: Suitable for scenarios requiring the model to engage in roleplay without refusing certain prompts.
  • Research: Useful for studying alignment, refusal mechanics, or the effects of targeted weight edits on LLM behavior.
  • Use Cases Blocked by RLHF: Addresses scenarios where the over-refusal common in RLHF-era models is a hindrance.