Full spec sheet, single-GPU performance, and multi-GPU scaling for the entire RTX 30 Series desktop lineup, as they behave under llama.cpp.
| Card | Die | CUDA cores | Boost | VRAM | Bus | Bandwidth | TDP | NVLink |
|---|---|---|---|---|---|---|---|---|
| RTX 3050 | GA107 | 2560 | 1.78 GHz | 8 GB GDDR6 | 128-bit | 224 GB/s | 130 W | — |
| RTX 3060 8G | GA106 | 3584 | 1.78 GHz | 8 GB GDDR6 | 128-bit | 256 GB/s | 170 W | — |
| RTX 3060 12G | GA106 | 3584 | 1.78 GHz | 12 GB GDDR6 | 192-bit | 360 GB/s | 170 W | — |
| RTX 3070 | GA104 | 5888 | 1.73 GHz | 8 GB GDDR6 | 256-bit | 448 GB/s | 220 W | — |
| RTX 3070 Ti | GA104 | 6144 | 1.77 GHz | 8 GB GDDR6X | 256-bit | 912 GB/s | 290 W | — |
| RTX 3080 10G | GA102 | 8704 | 1.71 GHz | 10 GB GDDR6X | 320-bit | 760 GB/s | 320 W | — |
| RTX 3080 12G | GA102 | 8960 | 1.71 GHz | 12 GB GDDR6X | 384-bit | 912 GB/s | 350 W | — |
| RTX 3080 Ti | GA102 | 10240 | 1.67 GHz | 12 GB GDDR6X | 384-bit | 960 GB/s | 350 W | — |
| RTX 3090 | GA102 | 10496 | 1.70 GHz | 24 GB GDDR6X | 384-bit | 936 GB/s | 350 W | 112.5 GB/s |
| RTX 3090 Ti | GA102 | 10752 | 1.86 GHz | 24 GB GDDR6X | 384-bit | 1008 GB/s | 450 W | 112.5 GB/s |
Green rows: the two cards this site recommends most often. 3080 10G shown at launch spec (21 Gbps refresh ≈ 832 GB/s); 3080 12G bandwidth is derived (384-bit × 19 Gbps). All cards are PCIe 4.0 x16; all are GA10x (Ampere) and fully supported by current llama.cpp CUDA builds.
| Card | 8B Q4 | 13B Q4 | 30B Q4 | Prompt proc (13B) |
|---|---|---|---|---|
| RTX 3050 | ~22 t/s | ~12 t/s (Q3) | — | ~200 t/s |
| RTX 3060 8G | ~25 t/s | ~13 t/s (Q3) | — | ~350 t/s |
| RTX 3060 12G | ~35 t/s | ~20 t/s | — | ~350 t/s |
| RTX 3070 | ~45 t/s | ~23 t/s (Q3) | — | ~500 t/s |
| RTX 3070 Ti | ~90 t/s | ~50 t/s (Q3) | — | ~900 t/s |
| RTX 3080 10G | ~60 t/s | ~30 t/s (tight ctx) | — | ~700 t/s |
| RTX 3080 12G | ~70 t/s | ~40 t/s | — | ~900 t/s |
| RTX 3080 Ti | ~74 t/s | ~42 t/s | — | ~900 t/s |
| RTX 3090 | ~75 t/s | ~40 t/s | ~15 t/s | ~1400 t/s |
| RTX 3090 Ti | ~80 t/s | ~44 t/s | ~16 t/s | ~1800 t/s |
"—" = doesn't fit at that quantization. 13B "Q3" rows: model doesn't fit at Q4_K_M, shown at Q3_K_M for reference. Estimates from typical community benchmarks, ~8K context.
llama.cpp splits transformer layers across cards (-sm layer). Token generation
speed ≈ (sum of bandwidths) minus an inter-GPU tax: ~10–15% with NVLink, ~25–35% over
PCIe 4.0. Prompt processing scales nearly linearly in all cases.
| Build | GPU power | Interconnect | VRAM | 13B Q4 | 30B Q4 | 70B Q4_K_M |
|---|---|---|---|---|---|---|
| 2× 3060 12G | 340 W | PCIe 4.0 | 24 GB | ~35 t/s | ~11 t/s | — |
| 3× 3060 12G | 510 W | PCIe 4.0 | 36 GB | ~46 t/s | 27B Q6: ~11 t/s | — |
| 2× 3070 | 440 W | PCIe 4.0 | 16 GB | ~40 t/s | — | — |
| 3× 3070 | 660 W | PCIe 4.0 | 24 GB | ~40 t/s | 27B Q4: ~16 t/s | — |
| 2× 3080 12G | 700 W | PCIe 4.0 | 24 GB | ~61 t/s | ~25 t/s | — |
| 2× 3080 Ti | 700 W | PCIe 4.0 | 24 GB | ~65 t/s | ~26 t/s | — |
| 3× 3080 Ti | 1050 W | PCIe 4.0 | 36 GB | ~78 t/s | 27B Q6: ~26 t/s | — |
| 4× 3080 Ti | 1400 W | PCIe 4.0 | 48 GB | ~100 t/s | 27B Q8: ~24 t/s | ~21 t/s |
| 3090 + 3080 Ti | 700 W | PCIe 4.0 | 36 GB | ~60 t/s | ~22 t/s | ~6 t/s (Q2) |
| 2× 3090 | 700 W | NVLink | 48 GB | ~70 t/s | ~30 t/s | ~16 t/s |
| 3090 + 3090 Ti | 800 W | NVLink | 48 GB | ~72 t/s | ~31 t/s | ~17 t/s |
| 2× 3090 Ti | 900 W | NVLink | 48 GB | ~78 t/s | ~33 t/s | ~18 t/s |
| 3× 3090 | 1050 W | NVLink + PCIe | 72 GB | — | 27B Q8: ~25 t/s | ~20 t/s (Q4), Q5 at 32K+ |
| 3× 3090 Ti | 1350 W | NVLink + PCIe | 72 GB | — | 27B Q8: ~27 t/s | ~22 t/s (Q4) |
| 4× 3060 12G | 680 W | PCIe 4.0 | 48 GB | ~60 t/s | ~20 t/s | ~10 t/s |
| 4× 3090 | 1400 W | 2× NVLink | 96 GB | — | 27B Q8: ~32 t/s | ~28 t/s (Q5) |
| 4× 3090 Ti | 1800 W | 2× NVLink | 96 GB | — | 27B Q8: ~34 t/s | ~27 t/s (Q4) · Q8: ~17 |
llama.cpp's -ngl N puts the last N transformer layers on the CPU and the
rest on the GPUs. The number to internalize: offloaded layers run at CPU speed
(~1–3 t/s for a 27B model on a modern 8-core), so the penalty is not proportional:
| Scenario | On CPU | Result |
|---|---|---|
| 27B Q4 in 24 GB: last ~4 layers on CPU to fit | ~6% | ~85% of full-GPU speed |
| 27B Q4 in 16 GB: ~half the layers offloaded | ~50% | ~3–5 t/s — a different machine |
| 70B Q4 in 24 GB: ~2/3 offloaded | ~67% | ~2–3 t/s — demo territory |
--n-cpu-moe moves MoE expert weights to CPU — cheap, because only a few experts are read per token. Great for MoE models in tight VRAM.-ot output=CPU parks the big final layer (~0.5–1 GB on 27B/70B) on CPU — a free VRAM win with negligible speed cost.llama-server -m 27b-q4_k_m.gguf -ngl 60 -sm layer -c 8192 --tensor-split 1,1 -ot output=CPU
27B Q4 across 2 GPUs, last 4 layers + output embedding on CPU
llama-server -m model-q4_k_m.gguf -ngl 99 -sm layer --tensor-split 1,1llama-server -m model.gguf -ngl 99 -sm layer --tensor-split 60,40llama-server -m model.gguf -ngl 99 -sm row --n-cpu-moe 0 # row split for very asymmetric pairs
Qwen-class 27–32B models (Qwen3-32B, Qwen2.5-32B) are the most common "one size up" target. GGUF sizes: Q4_K_M ≈ 19 GB · Q5_K_M ≈ 22 GB · Q6_K ≈ 25 GB · Q8_0 ≈ 33 GB — plus KV cache. That puts hard lines on the table: Q4 needs ~20 GB (i.e. a 24 GB card), Q5 needs ~23 GB, Q6 needs 36 GB+, Q8 wants 48 GB.
| Build | VRAM | 27B Q4_K_M | 27B Q5_K_M | 27B Q6_K | 27B Q8_0 |
|---|---|---|---|---|---|
| 1× 3050 / 3060 8G / 3070 / 3070 Ti | 8 GB | ❌ (partial offload ~2–6 t/s) | ❌ | ❌ | ❌ |
| 1× 3060 12G / 3080 12G / 3080 Ti | 12 GB | ❌ (partial offload ~5–7 t/s) | ❌ | ❌ | ❌ |
| 1× 3080 10G | 10 GB | ❌ (partial offload ~6–8 t/s) | ❌ | ❌ | ❌ |
| 2× 3060 12G | 24 GB | ✅ ~10–12 t/s | ⚠️ ~9 t/s, 4–8K ctx | ❌ | ❌ |
| 2× 3080 | 20 GB | ⚠️ ~18–20 t/s, tight ctx | ❌ | ❌ | ❌ |
| 2× 3080 12G | 24 GB | ✅ ~24–26 t/s | ⚠️ ~19–21 t/s, 4–8K ctx | ❌ | ❌ |
| 2× 3080 Ti | 24 GB | ✅ ~25–28 t/s | ⚠️ ~20–22 t/s, 4–8K ctx | ❌ | ❌ |
| 1× 3090 | 24 GB | ✅ ~13–16 t/s | ⚠️ ~12–14 t/s, 8K ctx | ❌ (~1 GB over — IQ6_XS fits) | ❌ |
| 1× 3090 Ti | 24 GB | ✅ ~15–17 t/s | ⚠️ ~13–15 t/s, 8K ctx | ❌ (~1 GB over — IQ6_XS fits) | ❌ |
| 3090 + 3080 Ti | 36 GB | ✅ ~30 t/s | ✅ ~24 t/s | ✅ ~22 t/s | ⚠️ ~13–14 t/s, 8K ctx |
| 2× 3090 (NVLink) | 48 GB | ✅ (overkill) | ✅ ~24–27 t/s | ✅ ~22–25 t/s | ✅ ~18–21 t/s |
| 2× 3090 Ti (NVLink) | 48 GB | ✅ (overkill) | ✅ ~26–29 t/s | ✅ ~24–27 t/s | ✅ ~19–22 t/s |
| 3× 3060 12G | 36 GB | ✅ ~14 t/s | ✅ ~16 t/s | ✅ ~10–12 t/s | ❌ |
| 4× 3060 12G | 48 GB | ✅ (overkill) | ✅ ~14 t/s | ✅ ~12 t/s | ✅ ~9 t/s |
| 3× 3080 Ti | 36 GB | ✅ ~33 t/s | ✅ ~28 t/s | ✅ ~26 t/s | ⚠️ ~19 t/s, 8K ctx |
| 3× 3090 | 72 GB | ✅ (overkill) | ✅ (overkill) | ✅ ~27 t/s | ✅ ~25 t/s, 32K ctx |
| 4× 3090 (NVLink) | 96 GB | ✅ everything, long context: Q8 at ~30 t/s with 32K+ ctx | |||
| If you want… | Buy | Because |
|---|---|---|
| 8B models, minimum spend | 3060 8G / 3050 | Fast enough, cheap, cool |
| 13B at full quality, budget | 3060 12G | 12 GB for the price of an 8GB card |
| 13B, fast | 3080 Ti / 3080 12G | 960 / 912 GB/s + 12 GB — the Ti when prices are close, the 12G when it's cheaper |
| 27B (Qwen) at Q4 | 1× 3090 / 2× 3060 12G | 24 GB is the line: ~15 t/s single, ~11 t/s budget pair |
| 27B at Q6/Q8 | 2× 3090 | 48 GB + NVLink — Q6 ~23 t/s, Q8 ~20 t/s |
| Best all-round single card | 3090 | 24 GB, NVLink, 350 W — the reference card |
| Maximum single-card speed | 3090 Ti | 1008 GB/s, if the premium is small |
| 70B LLMs | 2× 3090 | 48 GB + NVLink = the 70B build |
| Budget big model (30B) | 2× 3060 12G | 24 GB, 340 W, ~11 t/s on 30B |