nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16

Hugging Face
TEXT GENERATIONPricing:Input $0.05 / Output $0.2Concurrent Unit Cost:2Model Size:30BQuant:FP8Context Size:32kPublished:Dec 4, 2025License:otherArchitecture:Transformer0.8K Warm

NVIDIA-Nemotron-3-Nano-30B-A3B-BF16 is a 30 billion parameter large language model developed by NVIDIA, featuring a hybrid Mixture-of-Experts (MoE) architecture with 3.5 billion active parameters. Designed for both reasoning and non-reasoning tasks, it can generate reasoning traces to improve solution quality, supporting a 1M token context length. This model excels in agentic reasoning, code, and general instruction following, and supports English, German, Spanish, French, Italian, and Japanese.

Loading preview...

Model Overview

NVIDIA-Nemotron-3-Nano-30B-A3B-BF16 is a 30 billion parameter large language model developed by NVIDIA, featuring a unique hybrid Mixture-of-Experts (MoE) architecture. It combines 23 Mamba-2 and MoE layers with 6 Attention layers, activating 6 experts per token from a total of 128 experts plus 1 shared expert per MoE layer, resulting in 3.5 billion active parameters. The model is designed for unified reasoning and non-reasoning tasks, capable of generating explicit reasoning traces for higher-quality solutions, which can be toggled via a chat template flag.

Key Capabilities

  • Advanced Reasoning: Excels in complex reasoning tasks, outperforming other models on benchmarks like AIME25 (with tools), MiniF2F, and SWE-Bench (OpenHands).
  • Hybrid MoE Architecture: Leverages a Mamba-2 and Transformer hybrid MoE design for efficient and accurate processing.
  • Multilingual Support: Supports English, German, Spanish, French, Italian, and Japanese, with improved performance using Qwen.
  • Long Context Handling: Capable of processing up to 1M tokens, demonstrated by strong performance on RULER-100 benchmarks at 256k, 512k, and 1M contexts.
  • Agentic Tasks: Shows strong performance in agentic benchmarks like SWE-Bench and TauBench V2.

Good For

  • AI Agent Systems: Ideal for developers building sophisticated AI agents that require robust reasoning capabilities.
  • Chatbots and RAG Systems: Suitable for creating advanced conversational AI and retrieval-augmented generation systems.
  • Instruction Following: Performs well on general instruction-following tasks, with configurable reasoning output.
  • Commercial Use: This model is ready for commercial deployment.