gemma-4-12b-it-GGUF

Try Try

gemma-4-12b-it-GGUF has 11.9B parameters (11.9B active per token) and takes 7.11 GB on disk. See where its parameters live, layer by layer, and what hardware it needs.

At a glance

unsloth/gemma-4-12b-it-GGUF
Repo createdMay 29, 2026updated Jul 17, 2026
Model typeDenseFull attention (GQA)
InputsText + imagesvision encoder in a separate mmproj file, not counted here
Total parameters11.9B
Active per token11.9B100% of the model
ExpertsNone (dense)
Max context (from config)256K tokens
Layers48
On disk7.11 GB1 file
PrecisionQ4_K (82%), Q6_K (18%)
QuantizationQ4_K_M (GGUF)mix of Q4_K, Q6_K · blocks of 256 · 100% of parameters
Fewest GPUs1× RTX 4060a single 8 GB card · weights only
Licenseapache-2.0
GitHubNot linked