Three series, side by side

Every desktop card the site tracks — RTX 30, 40 and 50 Series — on the two numbers that matter for llama.cpp: VRAM capacity and memory bandwidth, plus the multi-GPU interconnect story.

27 GPUs tracked 8–32 GB VRAM 224–1792 GB/s

Full specification sheet

CardSeriesDieVRAMBus BandwidthTDPNVLink
RTX 305030GA1078 GB GDDR6128-bit224 GB/s130 W—
RTX 3060 8G30GA1068 GB GDDR6128-bit256 GB/s170 W—
RTX 3060 12G30GA10612 GB GDDR6192-bit360 GB/s170 W—
RTX 307030GA1048 GB GDDR6256-bit448 GB/s220 W—
RTX 3070 Ti30GA1048 GB GDDR6X256-bit912 GB/s290 W—
RTX 3080 10G30GA10210 GB GDDR6X320-bit760 GB/s320 W—
RTX 3080 12G30GA10212 GB GDDR6X384-bit912 GB/s350 W—
RTX 3080 Ti30GA10212 GB GDDR6X384-bit960 GB/s350 W—
RTX 309030GA10224 GB GDDR6X384-bit936 GB/s350 W112.5 GB/s
RTX 3090 Ti30GA10224 GB GDDR6X384-bit1008 GB/s450 W112.5 GB/s
RTX 4050†40AD1078 GB GDDR6128-bit224 GB/s135 W—
RTX 4060 8G40AD1078 GB GDDR6128-bit272 GB/s115 W—
RTX 4060 Ti 8G40AD1068 GB GDDR6128-bit288 GB/s160 W—
RTX 4060 Ti 16G40AD10616 GB GDDR6128-bit288 GB/s160 W—
RTX 4070 12G40AD10312 GB GDDR6X192-bit504 GB/s200 W—
RTX 4070 Super 12G40AD10312 GB GDDR6X192-bit504 GB/s220 W—
RTX 4070 Ti 12G40AD10312 GB GDDR6X192-bit672 GB/s285 W—
RTX 4070 Ti Super 16G40AD10316 GB GDDR6X256-bit672 GB/s285 W—
RTX 4080 16G40AD10216 GB GDDR6X256-bit717 GB/s320 W—
RTX 4090 24G40AD10224 GB GDDR6X384-bit1008 GB/s450 W—
RTX 5050 8G50GB2078 GB GDDR6128-bit320 GB/s130 W—
RTX 5060 8G50GB2068 GB GDDR7128-bit448 GB/s145 W—
RTX 5060 Ti 16G50GB20616 GB GDDR7128-bit448 GB/s180 W—
RTX 5070 12G50GB20512 GB GDDR7192-bit672 GB/s250 W—
RTX 5070 Super 12G†50GB20512 GB GDDR7192-bit672 GB/s250 W—
RTX 5070 Ti 16G50GB20316 GB GDDR7256-bit896 GB/s300 W—
RTX 5080 16G50GB20316 GB GDDR7256-bit960 GB/s360 W—
RTX 5090 32G50GB20232 GB GDDR7512-bit1792 GB/s575 W—

Green rows: the site's most-recommended cards per tier. All cards are PCIe x16 (4.0 on the RTX 30/40 Series, 5.0 on the RTX 50 Series) and supported by current llama.cpp CUDA builds. RTX 30 Series card names link to full guides. † Not listed on NVIDIA’s official comparison page (fetched 2026-09-24) — specs on that card’s page are flagged as supplementary.

Memory bandwidth, all 28 cards

3050 / 4050
224 GB/s
5050
320 GB/s
4060
272 GB/s
4060 Ti 8G/16G
288 GB/s
3060 8G
256 GB/s
3060 12G
360 GB/s
3070 / 5060
448 GB/s
4070 / 4070 Super
504 GB/s
4070 Ti / 4070 Ti S / 5070 / 5070 S
672 GB/s
4080
717 GB/s
3080
760 GB/s
3080 12G
912 GB/s
5070 Ti
896 GB/s
3070 Ti
912 GB/s
3080 Ti
960 GB/s
3090
936 GB/s
5080
960 GB/s
3090 Ti / 4090
1008 GB/s
5090
1792 GB/s

On the same model, token-generation speed tracks this list almost linearly. The RTX 5090 alone covers the entire previous two generations' spread in one step above them.

VRAM tiers — what each step buys you

VRAM tierRTX 30 SeriesRTX 40 SeriesRTX 50 SeriesRuns (Q4_K_M + KV)
8 GB3050, 3060 8G, 3070, 3070 Ti4050, 4060, 4060 Ti 8G5050, 50608B Q8, 13B Q3–Q4 tight
10–12 GB3080 10G, 3080 12G, 3060 12G, 3080 Ti4070, 4070 S, 4070 Ti5070, 5070 S13B Q4 comfortable
16 GB—4060 Ti 16G, 4070 Ti S, 40805060 Ti 16G, 5070 Ti, 508013B Q8, MoE ~30B, 27B Q3–Q4
24 GB3090, 3090 Ti4090—27B Q4–Q5, 30B Q4, MoE ~30B
32 GB——509027B Q8, 30B Q5+, 70B Q2–IQ3
Reading the tiers Capacity steps matter more than bandwidth steps within a tier — an 8 GB 5060 (448 GB/s) runs exactly the same models as an 8 GB 3050 (224 GB/s), just ~2× faster. But a tier jump (12 → 16 GB, 24 → 32 GB) changes which models exist for you. Buy the tier first, then the bandwidth you can afford in it.

Multi-GPU across the series

llama.cpp splits transformer layers across cards (-sm layer). Token generation speed ≈ (sum of bandwidths) minus an inter-GPU tax:

InterconnectAvailable onLink speedSplit tax
NVLinkRTX 3090 / 3090 Ti only (2-way)112.5 GB/s~10–15%
PCIe 4.0 x16all other RTX 30 / 40 Series cards~31 GB/s per dir.~25–35%
PCIe 5.0 x16all RTX 50 Series cards~63 GB/s per dir.~20–30% (est.)

Full mechanics, CPU offloading and exact flags: multi-GPU guide →. RTX 30 Series build-by-build scaling tables: 30 Series comparison →.

Cross-series buyer's summary

If you want…BuyBecause
8B models, minimum spend3060 8G / 4060 / 5060All fine; used 3060 is cheapest per token
13B at full quality, budget3060 12G / 4070 / 507012 GB tier; pick by price
13B Q8 + MoE ~30B, single card4060 Ti 16G / 4070 Ti S / 5060 Ti 16GThe 16 GB tier is new since Ampere — the RTX 30 Series never had it
27B (Qwen) at Q4–Q53090 / 4090 / 2× 16 GB pair24 GB is the line (or 32 GB combined from two 16s)
27B at Q8, single card5090Only consumer card with 32 GB
70B LLMs, best value2× 3090 (NVLink)48 GB, ~10–15% split tax, cheapest $/GB of the big builds
70B LLMs, new hardware2× 5090 (PCIe 5.0)64 GB — Q4 with headroom, Q5/Q6 in long context
Maximum single-card speed50901792 GB/s — ~1.8× a 3090 Ti on token generation
Best used value309024 GB + NVLink at a used price; the reference card

Per-series comparisons

RTX 30 Series

Full spec sheet, single-GPU and multi-GPU tables for all 10 cards.
compare →

RTX 40 Series soon

Comparison tables land when the per-card guides do.
in progress

RTX 50 Series soon

Comparison tables land when the per-card guides do.
in progress
← home RTX 30 Series →