gemma-4-26B-A4B-it-GGUF

Try Try

gemma-4-26B-A4B-it-GGUF has 25.2B parameters (3.82B active per token) and takes 17.0 GB on disk. See where its parameters live, layer by layer, and what hardware it needs.

At a glance

unsloth/gemma-4-26B-A4B-it-GGUF
Repo createdApr 1, 2026updated Jul 17, 2026
Model typeMixture of expertsFull attention (GQA)
InputsText + imagesvision encoder in a separate mmproj file, not counted here
Total parameters25.2B
Active per token3.82B15% of the model
Experts8 of 128 active
Max context (from config)256K tokens
Layers30
On disk17.0 GB1 file
PrecisionQ4_K (49%), Q5_1 (32%), Q8_0 (16%), Q5_K (2.1%), other (0.27%)
QuantizationUD-Q4_K_XL (GGUF)mix of Q4_K, Q5_1, Q8_0 +1 more · blocks of 256 / blocks of 32 · 100% of parameters
Fewest GPUs1× RTX 4090a single 24 GB card · weights only
FamilyQuantization of google/gemma-4-26B-A4B-it
Licenseapache-2.0
GitHubNot linked