saidutta69/Qwen2.5-1.5B-Instruct-heretic
saidutta69/Qwen2.5-1.5B-Instruct-heretic is a 1.5 billion parameter instruction-tuned causal language model, a decensored variant of Qwen/Qwen2.5-1.5B-Instruct. Developed by RACER IS OP using Heretic v1.2.0, it suppresses refusal behavior via targeted weight edits rather than fine-tuning, preserving the base model's knowledge and instruction-following. This model is optimized for direct answers, making it suitable for local agents, roleplay, and research into alignment mechanics where typical LLM refusals are undesirable.
Loading preview...
Overview
saidutta69/Qwen2.5-1.5B-Instruct-heretic is a 1.5 billion parameter instruction-tuned model derived from Qwen/Qwen2.5-1.5B-Instruct. Its key differentiator is the deliberate suppression of refusal behavior, achieved through a process called "abliteration" using Heretic v1.2.0. This method involves targeted weight edits to the attention output and MLP down-projections, ensuring the base model's core knowledge and instruction-following capabilities remain largely intact.
Key Capabilities & Performance
- Decensored Responses: The model is engineered to provide direct answers, significantly reducing refusal rates. Performance metrics show refusals dropped from 99/100 to 1/100 adversarial prompts compared to the original model.
- Preserved Base Model Integrity: The abliteration process results in a low KL divergence of 0.1607 on the output distribution, indicating that the modification is narrow and targeted, not a broad perturbation of the model's original capabilities.
- Efficient Size: At 1.5 billion parameters, it is designed to run comfortably on CPU or low-VRAM GPUs, offering improved reasoning and coherence over smaller 0.5B variants.
Ideal Use Cases
- Local Agents & Roleplay: Suitable for applications requiring models that respond directly without typical RLHF-era over-refusal.
- Alignment Research: Valuable for researchers studying alignment, refusal mechanics, and methods of controlling model behavior.
- Overcoming Refusal Blocks: Useful for any scenario where the default refusal behavior of instruction-tuned models hinders desired functionality.