Qwen3.8-Flash-Next-FP8

Try Try

Qwen3.8-Flash-Next-FP8 has 180B parameters and takes 186 GB on disk. See where its parameters live, layer by layer, and what hardware it needs.

At a glance

Qwen/Qwen3.8-Flash-Next-FP8
Repo createdAug 24, 2026updated Aug 31, 2026
Model typeNot shown: breakdown incomplete
InputsText + images449M vision encoder
Total parameters180B
Active per tokenNot shown: breakdown incomplete
ExpertsNot shown: breakdown incomplete
Max context (from config)256K tokens
Layers4836 linear + 12 full attention
On disk186 GB131 files
PrecisionFP8 E4M3 (94%), BF16 (5.9%)
QuantizationFP8 E4M3blocks of 128×128 · 97% of parameters · attention, embeddings & output head, vision encoder and shared experts kept in BF16
Fewest GPUs1× NVIDIA B300a single 288 GB card · weights only
FamilyQuantization of Qwen/Qwen3.8-Flash-Next
Licenseqwen-community-1.0 (custom)
GitHubQwenLM/Qwen3.8-Flash-Next