Llama-3.1-8B-Instruct-4bit

Try Try

Llama-3.1-8B-Instruct-4bit has 8.03B parameters (7.50B active per token) and takes 4.52 GB on disk. See where its parameters live, layer by layer, and what hardware it needs.

At a glance

mlx-community/Llama-3.1-8B-Instruct-4bit
Repo createdFeb 15, 2025updated Feb 15, 2025
Model typeDenseFull attention (GQA)
InputsText
Total parameters8.03B
Active per token7.50B93% of the model
ExpertsNone (dense)
Max context (from config)128K tokens
Layers32
On disk4.52 GB1 file
Precision4-bit MLX (89%), scales (11%)
Quantization4-bit MLXgroups of 64 · 100% of parameters
Fewest GPUs1× RTX 4060a single 8 GB card · weights only
Licensellama3.1
GitHubNot linked