Best AI models for an 8 GB GPU (RTX 4060, RTX 5060)
Eight gigabytes is the most common amount of video memory, and it's enough for a real local assistant: models of up to about 8 billion parameters fit once they're compressed to 4 or 5 bits, with room for a useful amount of context.
The table lists the most downloaded models that fit on an 8 GB card with at least 8K tokens of context. It's a popularity ranking of what fits, not a quality ranking. Each name opens the full answer for an RTX 4060.
Nothing the site has analyzed fits with these settings.
0 models the site has analyzed fit on an RTX 4060 with at least 8K tokens of context. See them all in the hardware calculator.
What fits in 8 GB
About 90% of the card is usable, so roughly 7.2 GB. A 4-bit file takes a little over half a gigabyte per billion parameters, so:
- Up to about 4B parameters fits with plenty of room, even at 8-bit.
- 7–9B parameters fits at 4 or 5 bits (Q4_K_M, Q5_K_M), with a few thousand to tens of thousands of tokens of context depending on the model.
- Bigger models only fit below 3 bits per weight, where quality drops noticeably.
Tips
- Context costs memory. If a model loads but runs out of memory in long chats, lower the context length or switch the context cache to 8-bit in your app's settings.
- Smaller is often better here. A 4B model at 8-bit can be a better experience than an 8B model squeezed to 3 bits.
- For a different card, try the hardware calculator: pick your GPU and it lists everything that fits.