NVIDIA GeForce RTX 30 Series

Ampere. Where the local-LLM era started — and where the most important card ever shipped for llama.cpp lives: the RTX 3090, with 24 GB and the only consumer NVLink that was ever offered.

10 cards · 8–24 GB GA107 → GA102 NVLink on 3090 / 3090 Ti

Pick your card

Each page covers the card's specs, what those specs mean for llama.cpp, what models fit, expected speeds, and how it behaves in multi-GPU builds.

What makes this series different

Buying in 2026 Most of the RTX 30 Series is used-market now. The 3090 remains the reference used LLM card (24 GB + NVLink at a reasonable price); the 3060 12 GB remains the reference budget card. Newer series cards trade the NVLink advantage for GDDR6X/GDDR7 bandwidth and better perf-per-watt — see the RTX 40 Series and RTX 50 Series pages.

Single vs. multiple cards

llama.cpp splits a model's transformer layers across GPUs (tensor split). Two rules of thumb:

Each card page has a Multi-GPU section with the exact llama-cli / llama-server flags for that card. The series comparison has scaling tables for common 2-GPU and 4-GPU builds, and the multi-GPU guide covers the mechanics.

← home next: RTX 40 Series →