saidutta69/Mistral-Nemo-Instruct-heretic

TEXT GENERATIONConcurrent Unit Cost:1Model Size:12BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jul 21, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

The saidutta69/Mistral-Nemo-Instruct-heretic is a 12 billion parameter decensored variant of the Mistral AI and NVIDIA co-developed Mistral-Nemo-Instruct-2407 model, featuring a 32768 token context length. This model suppresses refusal behavior through targeted weight edits using the Heretic v1.4.0 abliteration method, rather than fine-tuning, preserving the base model's strong reasoning and multilingual capabilities. It is designed for developers requiring a model that provides direct answers and reliable instruction-following without censorship for complex reasoning and long-context applications.

Loading preview...

Overview

saidutta69/Mistral-Nemo-Instruct-heretic is a 12 billion parameter language model, a decensored version of the mistralai/Mistral-Nemo-Instruct-2407 model, co-developed by Mistral AI and NVIDIA. It maintains the base model's strong reasoning and long-context capabilities (32768 tokens) while suppressing refusal behaviors.

Key Capabilities

  • Decensored Responses: Achieves uncensored output by using the Heretic v1.4.0 abliteration method, which involves targeted weight edits to attention and MLP down-projections.
  • Preserved Base Model Strengths: Unlike fine-tuning, abliteration leaves the core knowledge and capabilities of the original Mistral-Nemo-Instruct-2407 model largely intact, ensuring strong reasoning and multilingual support.
  • Reliable Instruction-Following: Designed to provide direct answers and follow instructions without the typical refusals found in RLHF'd models.

Why Abliteration?

Traditional fine-tuning to counter RLHF-induced refusals can degrade model coherence. Abliteration, as detailed in the original writeup, specifically targets and edits the weights responsible for refusal, preserving the network's other functions.

Good For

  • Developers needing a 12B multilingual model that answers directly.
  • Complex reasoning tasks requiring uninhibited responses.
  • Long-context applications where censorship might hinder utility.
  • Use cases demanding reliable instruction-following without built-in refusal behaviors.