saidutta69/Mistral-Nemo-Instruct-heretic

Hugging Face
TEXT GENERATIONPricing:Input $0.87 / Cached $0.2 / Output $0.99Concurrent Unit Cost:1Model Size:12BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jul 21, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Warm

Mistral-Nemo-Instruct-heretic is a 12 billion parameter language model developed by RACER IS OP, based on the Mistral AI and NVIDIA co-developed Mistral-Nemo-Instruct-2407. This model features a 32768 token context length and is specifically decensored using the Heretic v1.4.0 'abliteration' method, which suppresses refusal behavior via targeted weight edits rather than fine-tuning. It is designed for developers requiring a multilingual model that provides direct answers for complex reasoning and long-context applications without censorship.

Loading preview...

Model Overview

saidutta69/Mistral-Nemo-Instruct-heretic is a 12 billion parameter language model derived from the mistralai/Mistral-Nemo-Instruct-2407 base model, co-developed by Mistral AI and NVIDIA. This variant has been decensored using the Heretic v1.4.0 'abliteration' technique. Unlike traditional fine-tuning, abliteration directly edits specific weight directions responsible for refusal behaviors, preserving the base model's original knowledge and capabilities while eliminating censorship.

Key Capabilities and Features

  • Decensored Responses: Provides direct answers to queries that the base model might refuse, ensuring reliable instruction-following.
  • High Fidelity: Maintains the strong reasoning and long-context capabilities (32768 tokens) of the original Mistral-Nemo-Instruct-2407.
  • Multilingual Support: Inherits multilingual capabilities from its base model.
  • Efficient Deployment: Available with a full GGUF ladder, optimized for various GPU configurations from 6GB to 24GB, making it suitable for gaming PCs and CPU-only environments.
  • Targeted Modification: Utilizes directional ablation to suppress refusal behavior by editing attention output and MLP down-projections, avoiding the coherence degradation often seen with fine-tuning.

Ideal Use Cases

This model is particularly suited for developers who need:

  • A 12B multilingual model that provides direct, uncensored responses.
  • Reliable instruction-following for complex reasoning tasks.
  • Applications requiring long-context processing without model refusals.
  • Deployment on consumer-grade hardware, with various GGUF quantizations available.