Qwen3-235B-A22B

Try Try

Qwen3-235B-A22B has 235B parameters (21.6B active per token) and takes 470 GB on disk. See where its parameters live, layer by layer, and what hardware it needs.

At a glance

Qwen/Qwen3-235B-A22B
Repo createdApr 27, 2025updated Jul 26, 2025
Model typeMixture of expertsFull attention (GQA)
InputsText
Total parameters235B
Active per token21.6B9.2% of the model
Experts8 of 128 active
Max context (from config)40K tokens
Layers94
On disk470 GB118 files
PrecisionBF16 (100%)
QuantizationNoneoriginal precision (BF16)
Fewest GPUs7× NVIDIA H100fits in one 8-GPU server · weights only
FamilyOriginal, with derivatives · 3 models built on this
Licenseapache-2.0
GitHubQwenLM/Qwen3