saidutta69/Mistral-7B-Instruct-v0.3-heretic
saidutta69/Mistral-7B-Instruct-v0.3-heretic is a 7 billion parameter instruction-tuned causal language model, a decensored variant of Mistral-7B-Instruct-v0.3. Developed by RACER IS OP, it utilizes 'abliteration' to suppress refusal behavior via targeted weight edits, leaving the base model's knowledge and instruction-following capabilities largely intact. This model is optimized for use cases requiring an uncensored general-purpose model, such as local agents, roleplay, and function-calling, while running efficiently on consumer hardware.
Loading preview...
Model Overview
This model, Mistral-7B-Instruct-v0.3-heretic, is a 7 billion parameter instruction-tuned variant of the original mistralai/Mistral-7B-Instruct-v0.3. Developed by RACER IS OP, its primary distinction is the deliberate suppression of refusal behaviors through a technique called 'abliteration' (directional ablation) using Heretic v1.4.0. Unlike fine-tuning, abliteration directly edits specific weights responsible for refusals, aiming to preserve the base model's core knowledge and instruction-following abilities.
Key Capabilities & Features
- Decensored Output: Significantly reduces refusal rates (from 86/100 to 3/100 adversarial prompts) compared to the base model, enabling more direct responses.
- Preserved Base Capabilities: Maintains the original
Mistral-7B-Instruct-v0.3's instruction-following and general knowledge due to the targeted nature of abliteration. - Efficient for Consumer Hardware: Provided with a full suite of GGUF quantizations (Q8_0, Q6_K, Q5_K_M, Q4_K_M) to run on various GPUs and even CPU-only setups.
- Low KL Divergence: The abliteration process results in a low KL divergence (0.0687) from the base model, indicating minimal alteration to the overall output distribution.
Ideal Use Cases
- Local Agents: Suitable for applications where an uncensored model is preferred for autonomous agents.
- Roleplay & Creative Writing: Excels in scenarios requiring unrestricted creative output and character interaction.
- Function-Calling: Can be used for tasks where the model needs to comply with specific function calls without refusal.
- Research: Valuable for studying alignment, refusal mechanics, and the impact of targeted weight edits on LLM behavior.