Qwen3.8-27B-GGUF

Try Try

Qwen3.8-27B-GGUF has 27.3B parameters (25.6B active per token) and takes 13.1 GB on disk. See where its parameters live, layer by layer, and what hardware it needs.

At a glance

byteshape/Qwen3.8-27B-GGUF
Repo createdAug 18, 2026updated Sep 18, 2026
Model typeDenseHybrid attention: 48 linear + 17 full (GQA)
InputsText + imagesvision encoder in a separate mmproj file, not counted here
Total parameters27.3B
Active per token25.6B94% of the model
ExpertsNone (dense)
Max context (from config)256K tokens
Layers6448 linear + 17 full attention · plus 1 extra prediction layer
On disk13.1 GB1 file
PrecisionIQ4_XS (28%), IQ3_S (27%), IQ3_XXS (19%), Q4_K (13%), Q6_K (8.2%), Q3_K (2.9%), Q5_K (1.6%), other (0.96%)
QuantizationIQ4_XS (GGUF)mix of IQ3_S, IQ4_XS, IQ3_XXS +4 more · blocks of 256 · 100% of parameters
Fewest GPUs1× RTX 5060 Ti 16 GBa single 16 GB card · weights only
Licenseapache-2.0
GitHubNot linked