Qwen3.8-Flash-Next-GGUF

Try Try

Qwen3.8-Flash-Next-GGUF has 3.88B parameters (636M active per token) and takes 2.78 GB on disk. See where its parameters live, layer by layer, and what hardware it needs.

At a glance

unsloth/Qwen3.8-Flash-Next-GGUF
Repo createdAug 26, 2026updated Sep 2, 2026
Model typeMixture of expertsFull attention (GQA)
InputsText + imagesvision encoder in a separate mmproj file, not counted here
Total parameters3.88B
Active per token636M16% of the model
Experts10 of 512 activeplus 1 shared, always on
Max context (from config)256K tokens
Layers48the files show 0, which doesn't match; layer details hidden
On disk2.78 GB1 file
PrecisionQ4_K (48%), Q8_0 (32%), Q6_K (19%), other (0.56%)
QuantizationQ4_K_M (GGUF)mix of Q4_K, Q8_0, Q6_K · blocks of 256 / blocks of 32 · 100% of parameters
Fewest GPUs1× RTX 4060a single 8 GB card · weights only
Licenseqwen-community-1.0 (custom)
GitHubNot linked