Which AI models can I run on an RTX 4080?

0 of the 0 LLMs analyzed on Model Anatomy fit on an RTX 4080 (14.4 GB usable) with at least 4K tokens of context. Each row shows the best file that fits and how much context the remaining memory holds.

Nothing the site has analyzed fits with these settings. Try less context, an 8-bit context cache or more GPUs.

Assumes 90% of each card's memory is usable. Context memory is for one conversation. GGUF repos use their largest file that fits; other repos use the published weights, or a 4-bit estimate when only that would fit.