trohrbaugh/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16-heretic

TEXT GENERATIONConcurrent Unit Cost:2Model Size:30BQuant:FP8Context Size:32kPublished:Mar 18, 2026License:nvidia-nemotron-open-model-licenseArchitecture:Transformer0.0K Open Weights Featherless Exclusive Cold

The trohrbaugh/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16-heretic is a 30 billion parameter decensored version of NVIDIA's Nemotron-3-Nano-30B-A3B-BF16 model, created using the Heretic v1.2.0 tool with Arbitrary-Rank Ablation. This hybrid Mixture-of-Experts (MoE) model, combining Mamba-2 and Attention layers, is designed for both reasoning and non-reasoning tasks, supporting a 32K context length with a maximum of 1M tokens. It significantly reduces refusals compared to the original model, making it suitable for AI agent systems, chatbots, and RAG applications requiring less restrictive content generation.

Loading preview...

Model Overview

This model, trohrbaugh/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16-heretic, is a decensored variant of the NVIDIA Nemotron-3-Nano-30B-A3B-BF16, processed with the Heretic v1.2.0 tool using Arbitrary-Rank Ablation (ARA). The original NVIDIA Nemotron-3-Nano-30B-A3B-BF16 is a 30 billion parameter hybrid Mixture-of-Experts (MoE) model, developed by NVIDIA, integrating Mamba-2 and Attention layers. It is designed for both reasoning and non-reasoning tasks, capable of generating reasoning traces before a final response, which generally improves solution quality.

Key Differentiators

  • Decensored Behavior: Significantly reduces refusals (6/100 vs. 99/100 for the original model), offering less restrictive content generation.
  • Hybrid MoE Architecture: Combines 23 Mamba-2 and MoE layers with 6 Attention layers, featuring 128 experts plus 1 shared expert per MoE layer.
  • Reasoning Capabilities: Can be configured to generate explicit reasoning traces for complex tasks, enhancing accuracy, or provide direct answers for faster inference.
  • Multilingual Support: Supports English, German, Spanish, French, Italian, and Japanese, along with 43 programming languages.
  • Extended Context Window: Supports a maximum input/output context of 1 million tokens.

Ideal Use Cases

  • AI Agent Systems: Building intelligent agents that require robust reasoning and less content restriction.
  • Chatbots and Conversational AI: Developing chatbots that can handle diverse queries without excessive filtering.
  • RAG Systems: Enhancing Retrieval-Augmented Generation applications with a model capable of detailed reasoning.
  • Instruction Following: General instruction-following tasks where a more open-ended response is desired.