nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-Base-BF16

TEXT GENERATIONPricing:Input $0.2 / Output $0.8Concurrent Unit Cost:2Model Size:30BQuant:FP8Context Size:32kPublished:Dec 3, 2025License:otherArchitecture:Transformer0.1K Featherless Exclusive Cold

NVIDIA-Nemotron-3-Nano-30B-A3B-Base-BF16 is a 30 billion parameter base large language model developed by NVIDIA, featuring a Mamba2-Transformer Hybrid Mixture of Experts (MoE) architecture. It is pre-trained for next token prediction on a vast corpus of text, code, math, and science data, and supports a 32K context length. This model excels in mathematical reasoning, code generation, and long-context understanding, making it a strong foundation for building specialized instruction-following LLMs.

Loading preview...

Model Overview

NVIDIA-Nemotron-3-Nano-30B-A3B-Base-BF16 is a 30 billion parameter base large language model from NVIDIA, part of the Nemotron family of open models. It utilizes a Mamba2-Transformer Hybrid Mixture of Experts (MoE) architecture and is pre-trained for next token prediction. The model is designed for commercial use and serves as a robust starting point for instruction fine-tuning.

Key Capabilities & Performance

This model demonstrates strong performance across various benchmarks, often outperforming Qwen3 30B-A3B-Base in key areas:

  • Mathematical Reasoning: Achieves 92.34% on GSM8K and 82.88% on MATH, significantly higher than its comparator.
  • Code Generation: Scores 78.05% on HumanEval and 75.49% on MBPP-Sanitized.
  • Long Context Understanding: Supports up to 512K tokens, with strong RULER scores (e.g., 87.50% at 64K and 70.56% at 512K), a notable differentiator.
  • Multilingual Support: Trained on 20 languages, including English, Spanish, French, German, Japanese, and Chinese, and 43 programming languages.

Training & Data

The model was pre-trained on over 13 trillion tokens, including a significant portion of high-quality curated and synthetically-generated data spanning code, math, science, and general knowledge. The training data has a cutoff date of June 25, 2025.

Intended Use

This model is intended for developers and researchers who are building instruction-following LLMs, particularly those requiring strong performance in mathematical, coding, and long-context tasks. It is optimized for NVIDIA GPU-accelerated systems.