Best AI models for a 24 GB GPU (RTX 4090, RTX 3090)

A 24 GB card like the RTX 4090, RTX 3090 or Radeon RX 7900 XTX is the sweet spot for running AI models at home. It holds most models up to about 30 billion parameters once they're compressed to 4 or 5 bits, with room left over for a long conversation.

The table below lists the most downloaded models that fit, with at least 16K tokens of context. "Best" here means popular and fits, not a quality ranking: the site measures what a model needs, not how well it answers. Each name opens a page with the full answer for an RTX 4090: which files fit, and how much context each leaves room for.

Nothing the site has analyzed fits with these settings.

0 models the site has analyzed fit on an RTX 4090 with at least 16K tokens of context. See them all in the hardware calculator.

How to read the table

Getting more out of 24 GB

To run a file, the simplest options are Ollama and LM Studio; each model's page has a copy-ready command.