nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16

Hugging Face
TEXT GENERATIONPricing:Input $0.125 / Cached $0.025 / Output $1.15Concurrent Unit Cost:4Model Size:120BQuant:FP8Context Size:256kPublished:Mar 10, 2026License:otherArchitecture:Transformer0.4K Warm

NVIDIA-Nemotron-3-Super-120B-A12B-BF16 is a 120 billion parameter (12B active) large language model developed by NVIDIA, featuring a LatentMoE architecture that combines Mamba-2, MoE, and Attention layers. Optimized for agentic workflows and long-context reasoning, it supports up to 1 million tokens and is designed for high-volume applications like IT ticket automation and tool use. The model incorporates Multi-Token Prediction (MTP) for faster generation and improved quality, and is trained with NVFP4 quantization for efficiency.

Loading preview...

Model Overview

NVIDIA Nemotron-3-Super-120B-A12B-BF16 is a 120 billion parameter (12B active) large language model from NVIDIA, part of the Nemotron family focused on open models for specialized AI agents. It features a unique LatentMoE architecture, blending Mamba-2, Mixture-of-Experts (MoE), and Attention layers, and is enhanced with Multi-Token Prediction (MTP) for accelerated and higher-quality text generation. The model supports an extensive 1 million token context length and is trained using NVFP4 quantization for computational efficiency.

Key Capabilities

  • Configurable Reasoning: The model can generate a reasoning trace before its final response, with this capability toggled via a chat template flag (enable_thinking=True/False). A low-effort reasoning mode is also available.
  • Multilingual Support: It handles English, French, German, Italian, Japanese, Spanish, and Chinese.
  • Agentic Workflows: Designed for collaborative agents, tool use, and complex instruction following.
  • High-Volume Workloads: Optimized for tasks such as IT ticket automation and RAG systems.
  • Speculative Decoding: Includes a built-in MTP head, with an updated MTPv2 head available for further performance gains.

Benchmarks & Performance

The model demonstrates competitive performance across various benchmarks, particularly excelling in long-context tasks, achieving 91.75 on RULER @ 1M. It also shows strong results in reasoning, agentic, and chat/instruction following categories, often outperforming or closely matching models like Qwen3.5-122B-A10B and GPT-OSS-120B in specific areas such as HMMT Feb25 (with tools) and SWE-Bench (OpenHands).

Good For

  • Developers building AI Agent systems requiring robust reasoning and tool-use capabilities.
  • Applications demanding long-context understanding and generation.
  • High-volume conversational AI and automation tasks.
  • Multilingual applications in the supported languages.

Popular Sampler Settings

Top 3 parameter combinations used by Featherless users for this model. Click a tab to see each config.

temperature
top_p
top_k
frequency_penalty
presence_penalty
repetition_penalty
min_p