saidutta69/Qwen2.5-7B-Instruct-heretic

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:7.6BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jul 16, 2026License:otherArchitecture:Transformer Featherless Exclusive Cold

The saidutta69/Qwen2.5-7B-Instruct-heretic is a 7.6 billion parameter instruction-tuned causal language model based on Qwen2.5, developed by saidutta69. This variant is decensored using Heretic v1.4.0's directional ablation, suppressing refusal behavior through targeted weight edits rather than fine-tuning. It is optimized for use cases requiring direct answers without refusal, such as local agents, roleplay, or research into alignment mechanics, while largely preserving the base model's knowledge and instruction-following capabilities.

Loading preview...

Overview

This model, saidutta69/Qwen2.5-7B-Instruct-heretic, is a 7.6 billion parameter instruction-tuned variant of the Qwen2.5-7B-Instruct base model. It has been decensored using the Heretic v1.4.0 "abliteration" technique, which involves targeted weight edits to the attention output and MLP down-projections. This method suppresses refusal behavior without extensive fine-tuning, aiming to keep the base model's core knowledge and instruction-following intact.

Key Characteristics

  • Decensored Behavior: Significantly reduces refusal rates, achieving 3 refusals out of 100 adversarial prompts compared to 98/100 for the base model.
  • Abliteration Method: Utilizes directional ablation to modify specific weight directions responsible for refusal, minimizing impact on overall model coherence and capabilities.
  • Low KL Divergence: Exhibits a KL divergence of 0.0765 from the base model, indicating a narrow and targeted modification rather than a broad perturbation.
  • Local Deployment: Designed to be runnable on consumer-grade GPUs, with various GGUF quantizations provided (e.g., Q8_0 for 24GB VRAM, Q4_K_M for 8GB VRAM).

Use Cases

This model is particularly suited for developers who need a mid-size, locally-runnable Qwen2.5 model that provides direct answers without refusing or lecturing. It is ideal for:

  • Local Agents: Deploying AI agents that require uninhibited responses.
  • Roleplay: Creating interactive roleplay scenarios where refusal behavior is undesirable.
  • Research: Investigating alignment, refusal mechanics, or other areas where controlled model behavior is crucial.

It's important to note that this model removes refusal-shaped guardrails but does not inherently upgrade general capabilities or add safety filtering. Users are responsible for its deployment and ensuring ethical use, as it will comply with requests the base model would typically refuse.