unsloth/Nemotron-3-Nano-30B-A3B-Base

TEXT GENERATIONConcurrent Unit Cost:2Model Size:30BQuant:FP8Context Size:32kPublished:Dec 15, 2025License:nvidia-open-model-licenseArchitecture:Transformer0.0K Open Weights Featherless Exclusive Cold

The NVIDIA-Nemotron-3-Nano-30B-A3B-Base-BF16 is a 30 billion parameter base large language model developed by NVIDIA, featuring a Mamba2-Transformer Hybrid Mixture of Experts (MoE) architecture. Pre-trained on over 13 trillion tokens with a data cutoff of June 2025, it excels in mathematical reasoning, code generation, and long-context understanding up to 512K tokens. This model is designed for developers and researchers building instruction-following LLMs and supports 20 human languages and 43 programming languages.

Loading preview...

Model Overview

The NVIDIA-Nemotron-3-Nano-30B-A3B-Base-BF16 is a 30 billion parameter base large language model developed by NVIDIA, utilizing a Mamba2-Transformer Hybrid Mixture of Experts (MoE) architecture. It is pre-trained from scratch with a next-token prediction loss, making it an excellent foundation for instruction fine-tuning and commercial applications. The model was trained on an extensive corpus of over 13 trillion tokens, including a significant portion of synthetic data, with a data freshness cutoff of June 2025.

Key Capabilities & Performance

  • Advanced Reasoning: Demonstrates strong performance in mathematical tasks, achieving 92.34% on GSM8K and 82.88% on MATH, significantly outperforming Qwen3 30B-A3B-Base.
  • Code Generation: Excels in coding benchmarks, scoring 78.05% on HumanEval and 75.49% on MBPP-Sanitized.
  • Long Context Understanding: Features exceptional long-context capabilities, with RULER scores of 87.50% at 64K tokens, 82.92% at 128K tokens, 75.44% at 256K tokens, and 70.56% at 512K tokens.
  • Multilingual Support: Supports 20 human languages (including English, Spanish, French, German, Japanese, Chinese, Korean, Russian) and 43 programming languages.
  • Hybrid MoE Architecture: Combines Mamba2 and Transformer elements in a Mixture of Experts design for enhanced efficiency and accuracy.

Ideal Use Cases

  • Instruction Fine-tuning: Serves as a robust base model for developers and researchers to create specialized instruction-following LLMs.
  • Code Development: Suitable for applications requiring high-accuracy code generation and understanding across multiple programming languages.
  • Complex Problem Solving: Recommended for tasks involving intricate mathematical reasoning and general knowledge requiring deep understanding.
  • Long Document Analysis: Excellent for processing and generating content from very long texts, leveraging its extensive context window.