Which AI models can I run on my GPU?

Pick your GPU, several GPUs or a Mac to see which of the 0 open-weight AI models (LLMs) analyzed on Model Anatomy it can run, which file to download, and how much context the remaining memory holds. The counts below are for one device and at least 4K tokens of context.

Data center

GeForce

Radeon

Workstation

Mac (Apple Silicon)

Usable memory: 90% of each GPU, and about two thirds (up to 36 GB) or three quarters of a Mac's memory, macOS's default limit for the GPU. Each device's page can add more GPUs, more context and an 8-bit context cache.