Inkling-Small-GGUF

Try Try

Inkling-Small-GGUF has 264B parameters (11.2B active per token) and takes 163 GB on disk. See where its parameters live, layer by layer, and what hardware it needs.

At a glance

unsloth/Inkling-Small-GGUF
Repo createdJul 30, 2026updated Jul 31, 2026
Model typeMixture of expertsFull attention (GQA)
InputsText + imagesvision encoder in a separate mmproj file, not counted here
Total parameters264B
Active per token11.2B4.2% of the model
Experts6 of 256 activeplus 2 shared, always on
Max context (from config)1M tokens
Layers42
On disk163 GB5 files
PrecisionQ4_K (58%), Q5_K (36%), Q8_0 (3.8%), Q6_K (2.2%), other (0.11%)
QuantizationUD-Q4_K_XL (GGUF)mix of Q4_K, Q5_K, Q8_0 +1 more · blocks of 256 / blocks of 32 · 100% of parameters
Fewest GPUs1× NVIDIA B300a single 288 GB card · weights only
Licenseapache-2.0
GitHubNot linked