Ishowbackup/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16-heretic

TEXT GENERATIONPricing:Input $0.2 / Output $0.8Concurrent Unit Cost:2Model Size:30BQuant:FP8Context Size:32kPublished:Aug 13, 2026License:otherArchitecture:Transformer Featherless Exclusive Cold

Ishowbackup/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16-heretic is a 30 billion parameter decensored version of NVIDIA's Nemotron-3-Nano-30B-A3B-BF16, created using the Heretic v1.2.0 tool with Arbitrary-Rank Ablation. This model features a hybrid Mixture-of-Experts (MoE) architecture, combining Mamba-2 and Attention layers, and is designed for both reasoning and non-reasoning tasks with configurable reasoning traces. It supports English, German, Spanish, French, Italian, and Japanese, and excels in agentic and long-context tasks, demonstrating significantly reduced refusals compared to its original counterpart.

Loading preview...

Model Overview

This model, Ishowbackup/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16-heretic, is a decensored variant of NVIDIA's Nemotron-3-Nano-30B-A3B-BF16, processed with the Heretic v1.2.0 tool using Arbitrary-Rank Ablation (ARA). The original NVIDIA model is a 30 billion parameter large language model (LLM) developed by NVIDIA Corporation, featuring a unique hybrid Mixture-of-Experts (MoE) architecture that integrates Mamba-2 and Attention layers. It is designed for unified reasoning and non-reasoning tasks, capable of generating reasoning traces before providing a final response, which can be configured for higher accuracy on complex prompts.

Key Differentiators

  • Decensored Version: Significantly reduces refusals from 99/100 in the original model to 6/100, offering broader utility.
  • Hybrid MoE Architecture: Combines 23 Mamba-2 and MoE layers with 6 Attention layers, including 128 experts plus 1 shared expert per MoE layer, with 6 experts activated per token.
  • Configurable Reasoning: Users can enable or disable intermediate reasoning traces, balancing between direct answers and higher accuracy for challenging tasks.
  • Extensive Multilingual Support: Supports English, German, Spanish, French, Italian, and Japanese, with improved performance using Qwen.
  • Long Context Window: Supports a maximum input and output size of 1 million tokens, though default Hugging Face configuration uses 256k tokens due to VRAM requirements.

Performance Highlights

  • Reduced Refusals: Achieves 6/100 refusals compared to 99/100 in the original model.
  • Strong Agentic Performance: Outperforms other models in benchmarks like AIME25 (with tools) at 99.2%, MiniF2F pass@1 at 50.0%, and SWE-Bench (OpenHands) at 38.8%.
  • Long Context Capability: Achieves high scores on RULER-100 benchmarks, including 92.9% at 256k and 91.3% at 512k context lengths.

Ideal Use Cases

  • AI Agent Systems: Well-suited for developing sophisticated AI agents that require robust reasoning capabilities.
  • Chatbots and RAG Systems: Effective for building advanced conversational AI and retrieval-augmented generation applications.
  • Instruction Following: Capable of handling typical instruction-following tasks across various domains.
  • Commercial Applications: Ready for commercial deployment under the NVIDIA Nemotron Open Model License.