minjaechoi/nemotron3-nano-30b-a3b-2p03bit-r22

TEXT GENERATIONPricing:Input $0.2 / Output $0.8Concurrent Unit Cost:2Model Size:30BQuant:FP8Context Size:32kPublished:Sep 20, 2026License:otherArchitecture:Transformer Featherless Exclusive Cold

The NVIDIA Nemotron-3-Nano-30B-A3B-BF16 is a 30 billion parameter large language model developed by NVIDIA, featuring a hybrid Mixture-of-Experts (MoE) architecture with 3.5B active parameters and a 1M token context length. Designed for both reasoning and non-reasoning tasks, it can generate reasoning traces to improve accuracy or provide direct answers. This model supports English, German, Spanish, French, Italian, and Japanese, and is optimized for commercial use in AI agent systems, chatbots, and RAG applications.

Loading preview...

Model Overview

NVIDIA Nemotron-3-Nano-30B-A3B-BF16 is a 30 billion parameter large language model (LLM) developed by NVIDIA, featuring a unique hybrid Mixture-of-Experts (MoE) architecture. It combines 23 Mamba-2 and MoE layers with 6 Attention layers, utilizing 3.5 billion active parameters. The model is designed for unified reasoning and non-reasoning tasks, capable of generating a reasoning trace before providing a final response, which can be configured for higher accuracy on complex prompts. It supports a substantial context length of up to 1 million tokens.

Key Capabilities

  • Advanced Reasoning: Can generate intermediate reasoning traces for improved accuracy on challenging tasks, with an option to disable this for direct answers.
  • Hybrid MoE Architecture: Employs a Mamba2-Transformer Hybrid Mixture of Experts design for efficient processing.
  • Multilingual Support: Supports English, German, Spanish, French, Italian, and Japanese, along with 43 programming languages.
  • Long Context Handling: Capable of processing up to 1 million tokens, though default Hugging Face configuration is 256k due to VRAM requirements.
  • Commercial Use Ready: Licensed under the NVIDIA Nemotron Open Model License for commercial applications.

Ideal Use Cases

  • AI Agent Systems: Designed for developers building intelligent AI agents.
  • Chatbots and Conversational AI: Suitable for general-purpose chat and instruction-following tasks.
  • RAG Systems: Can be integrated into Retrieval Augmented Generation (RAG) pipelines.
  • Code Generation: Trained on extensive code data, supporting 43 programming languages.