inkling-GGUF

Try Try

inkling-GGUF has 947B parameters (39.8B active per token) and takes 587 GB on disk. See where its parameters live, layer by layer, and what hardware it needs.

At a glance

unsloth/inkling-GGUF
Repo createdJul 14, 2026updated Jul 16, 2026
Model typeMixture of expertsFull attention (GQA)
InputsText + imagesvision encoder in a separate mmproj file, not counted here
Total parameters947B
Active per token39.8B4.2% of the model
Experts6 of 256 activeplus 2 shared, always on
Max context (from config)1M tokens
Layers66
On disk587 GB14 files
PrecisionQ4_K (58%), Q5_K (33%), Q6_K (4.7%), Q8_0 (3.5%)
QuantizationUD-Q4_K_XL (GGUF)mix of Q4_K, Q5_K, Q6_K +1 more · blocks of 256 / blocks of 32 · 100% of parameters
Fewest GPUs5× NVIDIA H200fits in one 8-GPU server · weights only
Licenseapache-2.0
GitHubNot linked