saidutta69/Qwen2.5-7B-Instruct-1M-heretic

TEXT GENERATIONConcurrent Unit Cost:1Model Size:7.6BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jul 21, 2026License:otherArchitecture:Transformer0.0K Featherless Exclusive Cold

Qwen2.5-7B-Instruct-1M-heretic is a 7.6 billion parameter instruction-tuned language model developed by RACER IS OP, based on Qwen/Qwen2.5-7B-Instruct-1M. This variant features a 1 million token context window and has its refusal behavior suppressed through targeted weight edits using the Heretic v1.4.0 abliteration method. It is designed for long-context uncensored reasoning, making it suitable for extensive document analysis and codebase-wide reasoning without censorship.

Loading preview...

Overview

Qwen2.5-7B-Instruct-1M-heretic is a 7.6 billion parameter instruction-tuned model derived from Qwen/Qwen2.5-7B-Instruct-1M, developed by RACER IS OP. Its key differentiator is the suppression of refusal behavior, achieved not through fine-tuning but via abliteration using Heretic v1.4.0. This method involves targeted weight edits to attention output and MLP down-projections, preserving the base model's original knowledge and capabilities while removing censorship.

Key Capabilities

  • Uncensored Reasoning: Delivers direct answers without refusals, even for requests the base model would typically decline.
  • Massive Context Window: Supports a 1 million token context, enabling deep analysis of very long documents or entire codebases.
  • Preserved Base Model Integrity: Abliteration ensures that the core knowledge and performance of the Qwen 2.5 base model remain largely intact.
  • GPU Accessibility: Optimized GGUF quants are provided, allowing the model to run efficiently on consumer GPUs with 8-12 GB VRAM (e.g., Q4_K_M on an RTX 4060).

Why Abliteration?

Traditional fine-tuning to remove refusals can degrade model coherence by fighting against the base model's original training. Abliteration, as detailed in the Heretic repo and the original writeup, directly edits the specific weight directions responsible for refusal, leaving other network functions untouched.

Use Cases

This model is ideal for developers requiring a 7B model with extensive context and direct, uncensored responses. It excels in scenarios such as:

  • Long-document analysis
  • Codebase-wide reasoning
  • Any application demanding massive context without built-in censorship or refusal behaviors.