glm-4-9b-chat-IMat-GGUF

Try Try

glm-4-9b-chat-IMat-GGUF has 9.40B parameters (8.78B active per token) and takes 5.75 GB on disk. See where its parameters live, layer by layer, and what hardware it needs.

At a glance

legraphista/glm-4-9b-chat-IMat-GGUF
Repo createdJun 20, 2024updated Jun 20, 2024
Model typeDenseFull attention (GQA)
InputsText
Total parameters9.40B
Active per token8.78B93% of the model
ExpertsNone (dense)
Max context (from config)128K tokens
Layers40
On disk5.75 GB1 file
PrecisionQ4_K (64%), Q5_0 (23%), Q6_K (8.9%), Q5_1 (3.7%)
QuantizationQ4_K_S (GGUF)mix of Q4_K, Q5_0, Q6_K +1 more · blocks of 256 / blocks of 32 · 100% of parameters
Fewest GPUs1× RTX 4060a single 8 GB card · weights only
Licenseglm-4 (custom)
GitHubNot linked