saidutta69/Qwen2.5-7B-Instruct-heretic
Qwen2.5-7B-Instruct-heretic by saidutta69 is a 7.6 billion parameter instruction-tuned causal language model based on Qwen2.5-7B-Instruct, featuring a 32K context length. This variant is specifically modified using the Heretic v1.4.0 'abliteration' method to suppress refusal behaviors, making it a 'decensored' model. It is optimized for direct answers without refusals, suitable for local agents, roleplay, and research into alignment mechanics where typical RLHF-induced guardrails are undesirable.
Loading preview...
Qwen2.5-7B-Instruct-heretic: Decensored Qwen2.5-7B
This model, developed by saidutta69, is a specialized variant of the Qwen/Qwen2.5-7B-Instruct base model. It leverages the Heretic v1.4.0 'abliteration' technique to suppress refusal behaviors, effectively creating a 'decensored' version of the original 7.6 billion parameter model. Unlike fine-tuning, abliteration directly edits specific weight directions responsible for refusal, preserving the base model's core knowledge and instruction-following capabilities.
Key Differentiators
- Refusal Suppression: Significantly reduces refusal rates (3/100 adversarial prompts compared to 98/100 for the base model) by targeting attention output and MLP down-projections.
- Capability Preservation: Aims to maintain the original Qwen2.5-7B-Instruct's knowledge and coherence, as the edits are narrow and targeted, resulting in a low KL divergence of 0.0765 from the base model's output distribution.
- Local Deployment Optimized: Full GGUF quantization ladder is provided, making it suitable for deployment on various consumer GPUs (e.g., RTX 3090, 4080, 3060) and CPU-only setups.
Ideal Use Cases
- Local Agents: For applications requiring direct, unfiltered responses without model-induced refusals.
- Roleplay: Where models need to adhere strictly to character without moralizing or declining prompts.
- Alignment Research: For studying refusal mechanics and model behavior when typical safety guardrails are removed.
- Development: For scenarios where RLHF-era over-refusal hinders specific application development.
It's important to note that this model is not a general capability upgrade; it is the Qwen2.5-7B-Instruct with its refusal mechanisms intentionally removed. Users are responsible for its deployment and ensuring appropriate use, as it lacks inherent safety filtering.