exolabs/nemotron-3-nano-30b-a3b-nvfp4-dequant-bf16
TEXT GENERATIONConcurrent Unit Cost:2Model Size:30BQuant:FP8Context Size:32kPublished:Jun 4, 2026License:otherArchitecture:Transformer Featherless Exclusive Cold
The exolabs/nemotron-3-nano-30b-a3b-nvfp4-dequant-bf16 model is a 30 billion parameter language model, derived from the Nemotron-3-Nano-30B-A3B-NVFP4 architecture. This specific version is a dequantized bfloat16 export, primarily intended for vLLM validation. It serves as a public checkpoint for evaluating the performance of the dequantized model, differing from the original upstream bfloat16 version.
Loading preview...
Model Overview
This model, exolabs/nemotron-3-nano-30b-a3b-nvfp4-dequant-bf16, is a 30 billion parameter language model. It is a public dequantized bfloat16 (BF16) export of the nemotron-3-nano-30b-a3b-nvfp4 internal model, specifically prepared for vLLM validation.
Key Characteristics
- Parameter Count: 30 billion parameters.
- Context Length: Supports a context length of 32768 tokens.
- Dequantized Export: This checkpoint is the result of dequantizing an MLX quantized checkpoint, meaning it is not the original upstream BF16 version.
- Export Dtype:
bfloat16. - Text-only: The export is not text-only, implying potential for other modalities if the base model supports them.
Intended Use
This model is primarily intended for:
- vLLM Validation: Its main purpose is to facilitate validation within the vLLM framework.
- Research and Development: Useful for researchers and developers interested in evaluating dequantized models and their performance characteristics compared to their quantized counterparts or original BF16 versions.