Can I run Qwen3.8-Flash-Next-GGUF on an RTX 5060 Ti 16 GB?
Yes
Its mtp-Qwen3.8-Flash-Next-BF16 file (7.77 GB) fits in the 14.4 GB an RTX 5060 Ti 16 GB can use, with room for an unknown amount of context.
Full analysis of Qwen3.8-Flash-Next-GGUF
Everything that runs on RTX 5060 Ti 16 GB
Get the mtp-Qwen3.8-Flash-Next-BF16 file on Hugging Face
ollama run hf.co/unsloth/Qwen3.8-Flash-Next-GGUFFiles and context on RTX 5060 Ti 16 GB
Each file of Qwen3.8-Flash-Next-GGUF (and of GGUF versions the site has analyzed), and how much context fits next to it, with the context cache in 16-bit or 8-bit.
| File | Size | On RTX 5060 Ti 16 GB | Context (16-bit) | Context (8-bit) |
|---|---|---|---|---|
| BF16 | 354 GB | Doesn't fit | — | — |
| Qwen3.8-Flash-Next-Q8_0 | 188 GB | Doesn't fit | — | — |
| UD-Q6_K_XL | 169 GB | Doesn't fit | — | — |
| UD-Q5_K_XL | 158 GB | Doesn't fit | — | — |
| UD-Q4_K_XL | 111 GB | Doesn't fit | — | — |
| UD-IQ4_XS | 93.7 GB | Doesn't fit | — | — |
| UD-Q3_K_XL | 90.0 GB | Doesn't fit | — | — |
| UD-IQ3_XXS | 82.0 GB | Doesn't fit | — | — |
| UD-Q2_K_XL | 78.9 GB | Doesn't fit | — | — |
| UD-IQ1_M | 74.5 GB | Doesn't fit | — | — |
| UD-IQ1_S | 72.5 GB | Doesn't fit | — | — |
| mtp-Qwen3.8-Flash-Next-BF16 | 7.77 GB | Fits | not known | not known |
| mtp-Qwen3.8-Flash-Next-shared-BF16 | 5.23 GB | Fits | not known | not known |
| Q8_0 | 4.14 GB | Fits | not known | not known |
| mtp-Qwen3.8-Flash-Next-shared-Q8_0 | 2.79 GB | Fits | not known | not known |
| Q4_K_M | 2.79 GB | Fits | not known | not known |
| mtp-Qwen3.8-Flash-Next-shared-Q4_K_M | 1.91 GB | Fits | not known | not known |