GLM-4.7-Flash

Try Try

GLM-4.7-Flash has 31.2B parameters (3.58B active per token) and takes 62.4 GB on disk. See where its parameters live, layer by layer, and what hardware it needs.

At a glance

zai-org/GLM-4.7-Flash
Repo createdJan 19, 2026updated Jan 29, 2026
Model typeMixture of expertsFull attention (MLA)
InputsText
Total parameters31.2B
Active per token3.58B11% of the model
Experts4 of 64 activeplus 1 shared, always on
Max context (from config)198K tokens
Layers47plus 1 extra prediction layer
On disk62.4 GB48 files
PrecisionBF16 (100%)
QuantizationNoneoriginal precision (BF16)
Fewest GPUs1× NVIDIA H100a single 80 GB card · weights only
Licensemit
GitHubNot linked