NVIDIA-Nemotron-3-Super-120B-A12B-BF16

Try Try

NVIDIA-Nemotron-3-Super-120B-A12B-BF16 has 124B parameters (13.0B active per token) and takes 247 GB on disk. See where its parameters live, layer by layer, and what hardware it needs.

At a glance

nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16
Repo createdMar 10, 2026updated Aug 25, 2026
Model typeMixture of expertsHybrid attention: 40 linear + 48 full (GQA)
InputsText
Total parameters124B
Active per token13.0B11% of the model
Experts22 of 512 activeplus 1 shared, always on
Max context (from config)256K tokens
Layers8840 linear + 48 full attention
On disk247 GB50 files
PrecisionBF16 (100%)
QuantizationNoneoriginal precision (BF16)
Fewest GPUs1× NVIDIA B300a single 288 GB card · weights only
Licensenvidia-nemotron-open-model-license (custom)
GitHubNot linked