saidutta69/Mistral-7B-Instruct-v0.3-heretic

TEXT GENERATIONConcurrent Unit Cost:1Model Size:7BQuant:FP8Context Size:4kTool Calling:SupportedPublished:Jul 16, 2026License:otherArchitecture:Transformer Featherless Exclusive Cold

The saidutta69/Mistral-7B-Instruct-v0.3-heretic is a 7 billion parameter instruction-tuned causal language model, a decensored variant of Mistral-7B-Instruct-v0.3. Developed by RACER IS OP using the Heretic v1.4.0 directional ablation method, it suppresses refusal behavior via targeted weight edits rather than fine-tuning. This model is optimized for use cases requiring an uncensored Mistral 7B architecture, such as local agents, roleplay, and function-calling, while largely retaining the base model's knowledge and instruction-following capabilities.

Loading preview...

Overview

This model, Mistral-7B-Instruct-v0.3-heretic, is a 7 billion parameter instruction-tuned language model derived from mistralai/Mistral-7B-Instruct-v0.3. Its primary distinction is the removal of refusal behaviors through a technique called abliteration (directional ablation) using Heretic v1.4.0. Unlike fine-tuning, abliteration directly edits specific weights responsible for refusal, preserving the base model's core knowledge and instruction-following abilities.

Key Capabilities & Differentiators

  • Decensored Output: Significantly reduces refusal rates from 86/100 to 3/100 on adversarial prompts compared to the base model.
  • Preserved Base Capabilities: Maintains the original Mistral-7B-Instruct-v0.3's knowledge and instruction-following due to the targeted nature of abliteration, resulting in a low KL divergence of 0.0687 from the base model.
  • Efficient Modification: Achieves decensoring without extensive fine-tuning, which can degrade coherence in other methods.

Ideal Use Cases

  • Local Agents: Suitable for applications requiring an uncensored model on consumer hardware.
  • Roleplay & Creative Generation: Excels in scenarios where refusal guardrails are undesirable.
  • Function-Calling: Provides direct responses without built-in safety filters.
  • Research: Useful for studying alignment and refusal mechanisms in LLMs.

This model is provided with GGUF quantizations (Q4_K_M, Q5_K_M, Q6_K, Q8_0) for efficient deployment with tools like llama.cpp and Ollama. Users are responsible for its deployment, as it will comply with requests the base model would refuse.