Iambackup/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16

TEXT GENERATIONConcurrent Unit Cost:2Model Size:30BQuant:FP8Context Size:32kPublished:Jul 8, 2026License:nvidia-nemotron-open-model-licenseArchitecture:Transformer Open Weights Featherless Exclusive Cold

NVIDIA-Nemotron-3-Nano-30B-A3B-BF16 is a 30 billion parameter large language model developed by NVIDIA, featuring a hybrid Mixture-of-Experts (MoE) architecture with Mamba-2 and Attention layers. Designed for both reasoning and non-reasoning tasks, it can generate intermediate reasoning traces for higher accuracy on complex prompts or provide direct answers. This model supports English and coding languages, with additional support for German, Spanish, French, Italian, and Japanese, making it suitable for AI agent systems, chatbots, and RAG applications.

Loading preview...

Model Overview

NVIDIA-Nemotron-3-Nano-30B-A3B-BF16 is a 30 billion parameter large language model (LLM) developed by NVIDIA, distinguished by its hybrid Mixture-of-Experts (MoE) architecture. It combines 23 Mamba-2 and MoE layers with 6 Attention layers, featuring 128 experts plus 1 shared expert per MoE layer, with 6 experts activated per token. The model is designed for unified reasoning and non-reasoning tasks, capable of generating explicit reasoning traces for enhanced accuracy on challenging prompts, a feature configurable via the chat template.

Key Capabilities

  • Advanced Reasoning: Excels in tasks requiring complex thought processes, with an option to generate reasoning traces for improved accuracy.
  • Hybrid MoE Architecture: Leverages a unique Mamba-2 and Transformer-based MoE design for efficient processing.
  • Multilingual Support: Supports English, German, Spanish, French, Italian, and Japanese, alongside 43 programming languages.
  • Extensive Training: Pre-trained on 25 trillion tokens, including a significant portion of synthetic data, and further fine-tuned with supervised and reinforcement learning.
  • Long Context: Supports a maximum input and output context of 1 million tokens.

Use Cases

  • AI Agent Systems: Ideal for developing sophisticated AI agents that require robust reasoning capabilities.
  • Chatbots and Conversational AI: Suitable for general-purpose chat and instruction-following applications.
  • RAG Systems: Can be integrated into Retrieval Augmented Generation (RAG) systems for enhanced knowledge retrieval and response generation.
  • Code Generation: Trained on a vast corpus of code, making it effective for programming-related tasks.