saidutta69/Mistral-7B-Instruct-v0.3-heretic

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:7BQuant:FP8Context Size:4kTool Calling:SupportedPublished:Jul 16, 2026License:otherArchitecture:Transformer0.0K Featherless Exclusive Cold

saidutta69/Mistral-7B-Instruct-v0.3-heretic is a 7 billion parameter instruction-tuned causal language model, a decensored variant of Mistral-7B-Instruct-v0.3. Developed by RACER IS OP, it utilizes 'abliteration' to suppress refusal behavior via targeted weight edits, leaving the base model's knowledge and instruction-following capabilities largely intact. This model is optimized for use cases requiring an uncensored general-purpose model, such as local agents, roleplay, and function-calling, while running efficiently on consumer hardware.

Loading preview...

Model Overview

This model, Mistral-7B-Instruct-v0.3-heretic, is a 7 billion parameter instruction-tuned variant of the original mistralai/Mistral-7B-Instruct-v0.3. Developed by RACER IS OP, its primary distinction is the deliberate suppression of refusal behaviors through a technique called 'abliteration' (directional ablation) using Heretic v1.4.0. Unlike fine-tuning, abliteration directly edits specific weights responsible for refusals, aiming to preserve the base model's core knowledge and instruction-following abilities.

Key Capabilities & Features

  • Decensored Output: Significantly reduces refusal rates (from 86/100 to 3/100 adversarial prompts) compared to the base model, enabling more direct responses.
  • Preserved Base Capabilities: Maintains the original Mistral-7B-Instruct-v0.3's instruction-following and general knowledge due to the targeted nature of abliteration.
  • Efficient for Consumer Hardware: Provided with a full suite of GGUF quantizations (Q8_0, Q6_K, Q5_K_M, Q4_K_M) to run on various GPUs and even CPU-only setups.
  • Low KL Divergence: The abliteration process results in a low KL divergence (0.0687) from the base model, indicating minimal alteration to the overall output distribution.

Ideal Use Cases

  • Local Agents: Suitable for applications where an uncensored model is preferred for autonomous agents.
  • Roleplay & Creative Writing: Excels in scenarios requiring unrestricted creative output and character interaction.
  • Function-Calling: Can be used for tasks where the model needs to comply with specific function calls without refusal.
  • Research: Valuable for studying alignment, refusal mechanics, and the impact of targeted weight edits on LLM behavior.