NVIDIA-Nemotron-3-Nano-30B-A3B-BF16

Try Try

NVIDIA-Nemotron-3-Nano-30B-A3B-BF16 has 31.6B parameters (3.58B active per token) and takes 63.2 GB on disk. See where its parameters live, layer by layer, and what hardware it needs.

At a glance

nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16
Repo createdDec 4, 2025updated Aug 24, 2026
Model typeMixture of expertsHybrid attention: 23 linear + 6 full (GQA)
InputsText
Total parameters31.6B
Active per token3.58B11% of the model
Experts6 of 128 activeplus 1 shared, always on
Max context (from config)256K tokens
Layers5223 linear + 6 full attention
On disk63.2 GB13 files
PrecisionBF16 (100%)
QuantizationNoneoriginal precision (BF16)
Fewest GPUs1× NVIDIA H100a single 80 GB card · weights only
Licensenvidia-nemotron-open-model-license (custom)
GitHubNot linked