How to Pick a GPU for Local LLM in 2026 — Buyer's Guide
VRAM is the gate, bandwidth is the speedometer, and price-per-token is the only metric that matters. Here's how to choose without overpaying.
By Mohamed Meguedmi · 11 min read
Key Takeaways
- VRAM is the gate, not the goal. If a quantized model plus its KV cache doesn't fit in VRAM, nothing else matters. Aim for VRAM ≥ 1.3× model file size.
- The used RTX 3090 (24 GB) is still the best VRAM-per-dollar buy in May 2026, at $650–$780 street. Nothing under $1,000 beats it for 30B-class models.
- The RTX 5090 (32 GB, 1,792 GB/s) is the new single-GPU endgame for 70B-class models at Q4. Street price $2,400–$3,100; prebuilt systems often cheaper than the card alone.
- Apple M4 Max / M3 Ultra remain the only sane path to 70B–120B at home if you value silence and unified memory over raw tokens/sec.
- Avoid the RTX 4070/5070 12 GB trap. 12 GB is a 2024 spec at a 2026 price. Either go 16 GB+ or buy used 24 GB.
The Only Four Numbers That Matter
Every credible GPU comparison for local inference comes down to four numbers: VRAM capacity, memory bandwidth, compute (TFLOPS for prefill), and street price. Everything else — CUDA cores, ray-tracing units, display outputs — is noise for LLM workloads.
VRAM determines which models you can load. Bandwidth determines how fast tokens come out during decode. Compute determines how fast long prompts are processed (prefill). Price determines whether the decision is rational. Use our cost calculator to convert these into price-per-million-tokens once you've shortlisted two or three cards.
The VRAM-to-Model-Size Rule
For a quantized GGUF or AWQ model, plan for: VRAM_needed ≈ model_file_size × 1.25 + (context_length × 0.0005 GB per token per layer). A practical shortcut: take the file size on Hugging Face, multiply by 1.3, and that's your floor. A 32B model at Q4_K_M is ~19 GB on disk, so you need ~25 GB of VRAM headroom for a 32K context — which is why the 24 GB RTX 3090 is on the edge, and the 32 GB RTX 5090 is comfortable.
The 2026 GPU Tier List for Local LLMs
Prices below are May 2026 US street averages (eBay sold listings, Newegg, Micro Center). Tokens/sec figures are for Qwen3-Coder 32B at Q4_K_M with a 4K prompt, using llama.cpp b5421 unless noted. See our methodology page for the full bench protocol.
| GPU | VRAM | Bandwidth | Street price (USD) | 32B Q4 tok/s | $/GB VRAM |
|---|---|---|---|---|---|
| RTX 3060 12 GB | 12 GB | 360 GB/s | $210 | n/a (won't fit) | $17.5 |
| RTX 4060 Ti 16 GB | 16 GB | 288 GB/s | $340 | n/a (tight, 14B max) | $21.3 |
| RTX 3090 (used) | 24 GB | 936 GB/s | $720 | 34 | $30 |
| RTX 4090 | 24 GB | 1,008 GB/s | $1,750 | 42 | $73 |
| RTX 5080 Super 24 GB | 24 GB | 1,152 GB/s | $1,400 | 46 | $58 |
| RTX 5090 | 32 GB | 1,792 GB/s | $2,650 | 78 | $83 |
| RTX PRO 6000 Blackwell | 96 GB | 1,792 GB/s | $8,200 | 74 | $85 |
| Apple M4 Max 128 GB | ~96 GB usable | 546 GB/s | $4,700 (laptop) | 26 | $49 |
| Apple M3 Ultra 256 GB | ~192 GB usable | 819 GB/s | $7,400 (Studio) | 38 | $39 |
Match the GPU to the Model You'll Actually Run
Buying a GPU before knowing what you'll run is how people end up with a 12 GB card and regret. Decide the target model first.
If you live in 7B–14B territory (coding assistants, RAG)
A 16 GB card is the floor. The RTX 4060 Ti 16 GB at $340 is the rational pick if you only ever run Qwen3 8B or Llama 3.3 8B. Don't pay more unless you're sure you'll scale up. The 4060 Ti's 288 GB/s bandwidth is the bottleneck, not VRAM — expect 38–45 tok/s on 8B Q4.
If you want 30B-class models (Qwen3-Coder 32B, GLM-4.5-Air 32B)
This is the sweet spot of 2026, and it's where 24 GB cards earn their keep. The used RTX 3090 at ~$720 is the unbeatable value play. The RTX 5080 Super 24 GB at $1,400 is the new-warranty option, with ~35% more bandwidth than the 3090 and Blackwell's improved FP4 paths. Skip the RTX 4090 in 2026 — it's been outflanked from below (3090 used) and above (5080 Super new).
If you want 70B-class models (Llama 3.3 70B, Qwen3 72B)
You need 48 GB+ of VRAM at Q4, or you need Apple Silicon. Three credible paths:
- 2× RTX 3090 NVLink-less at ~$1,500 total. Tensor parallelism via vLLM or llama.cpp's
--split-mode row. Hot, loud, 700W draw, but it works. - RTX 5090 32 GB with aggressive quantization (Q3_K_M for 70B at ~28 GB). Acceptable quality for code and chat, single card, clean.
- Apple M3 Ultra 256 GB. Unified memory means 70B Q8 fits trivially, and even 120B models load. Slower than dual 3090s on prefill, but silent and idles at 8 W.
If you're building for a small team or business
The RTX PRO 6000 Blackwell 96 GB at $8,200 is the only single-card path to 70B at Q8 or 120B at Q4 with serious context. ECC memory, blower-style cooling, and a 600 W TDP that fits in a tower case. Compare against renting an H100 hour at $2.50/hr — break-even is roughly 3,300 hours, or about 18 months of weekday business use.
Bandwidth Is the Decode Speedometer
For batch-size-1 inference (the local case), token generation is memory-bandwidth bound. Kipp Lehmann's transformer arithmetic shows the ceiling: tok/s ≈ bandwidth / (2 × active_params_in_bytes). For a 32B model at Q4 (~16 GB of active weights touched per token):
- RTX 3090 at 936 GB/s → theoretical max ~29 tok/s. Measured: 34 tok/s (KV cache helps).
- RTX 5090 at 1,792 GB/s → theoretical max ~56 tok/s. Measured: 78 tok/s with Flash Attention 3 + FP4.
- M4 Max at 546 GB/s → theoretical max ~17 tok/s. Measured: 26 tok/s with MLX.
The lesson: bandwidth scales tok/s almost linearly within an architecture. When two cards have the same VRAM, buy the higher bandwidth one — full stop.
The Used Market Is Where the Wins Live
In May 2026, the used market is the rational buyer's market. Used RTX 3090 prices have stabilized around $700–$780 after two years of decline, because demand for the only sub-$1,000 24 GB card finally cleared the post-mining glut. Used RTX 4090s sit at $1,400–$1,650, with risk: many were mining or 24/7 inference cards.
The team rule: if a used 3090 is within 15% of the price of a new 4060 Ti 16 GB, buy the 3090. VRAM headroom beats warranty for a hobbyist.
Test protocol for any used card: run nvidia-smi -q -d POWER,TEMPERATURE under sustained Ollama load for 20 minutes. Hot spot > 95 °C or fan curves that ramp aggressively at idle = repad needed (or walk away).
What About AMD, Intel, and the Edge Cases?
AMD Radeon RX 7900 XTX / RX 9070 XT
The RX 7900 XTX 24 GB at $730 is tempting on paper. ROCm 6.4 on Linux finally delivers ~80% of CUDA performance in llama.cpp Vulkan and HIP backends. The catch: tooling fragility, no CUDA-only frameworks (no bitsandbytes-NF4, no exllamav2 v0.3+), and Windows support that still trails. Recommend only if you already run Linux and like debugging.
Intel Arc B580 24 GB "Battlemage Pro"
Released March 2026 at $499. Twenty-four gigabytes of VRAM at half the price of a used 3090 sounds revolutionary, but bandwidth is only 456 GB/s and IPEX-LLM is still catching up to llama.cpp's CUDA path. Measured 18 tok/s on 32B Q4. Watch this space for 12 months; don't bet a build on it yet.
Apple Silicon is the silent dark horse
For developers who already own a Mac, the M4 Pro 48 GB at $2,400 runs Qwen3 32B at 22 tok/s with zero fan noise and 28 W power draw. The unified memory architecture means no host-to-device copies — context windows scale without OOM panic. The downside: prefill is slow (a 16K-token prompt takes 9–14 seconds on M4 Max vs ~1.5 s on a 5090). For RAG and long-context retrieval, that latency matters.
Power, Noise, and the Hidden Costs
The RTX 5090 pulls a sustained 525 W during inference. At a US average $0.17/kWh, running it 8 hours a day for a year costs ~$260 — not nothing, and irrelevant compared to renting. But two 3090s at 700 W combined are loud and hot enough to require dedicated airflow. Budget $80–$150 for a quality 1000 W+ Platinum PSU and an additional intake fan.
Noise matters more than people admit. The blower-style RTX PRO 6000 measures 52 dB(A) at load — a vacuum cleaner in your office. The triple-fan open-shroud RTX 5090 measures 38 dB(A). The M-series Mac Studio: 22 dB(A). Pick accordingly if the machine sits next to your desk.
The BestLLMfor Verdict — May 2026
| Budget | Pick | Why |
|---|---|---|
| < $400 | RTX 4060 Ti 16 GB (new) or RTX 3060 12 GB (used $180) | Caps you at 8B–14B. Honest about the limit. |
| $700–$800 | Used RTX 3090 24 GB | Best VRAM-per-dollar in 2026. Period. |
| $1,400 | RTX 5080 Super 24 GB | New warranty, Blackwell FP4, near-3090×1.4 speed. |
| $2,500–$3,000 | RTX 5090 32 GB | Single-card 70B-Q3 or 32B-Q8. The 2026 endgame. |
| $5,000+ | Apple M3 Ultra 256 GB or 2× RTX 5090 | Mac if you value silence and 120B models. Dual 5090 if you need prefill speed. |
| $8,000+ | RTX PRO 6000 Blackwell 96 GB | Small-team server. ECC, 70B Q8, 120B Q4 single card. |
For ongoing benchmarks across these cards, the BestLLMfor public benchmark API (CC BY 4.0) and the open-source quelllm-mcp server expose the raw tok/s figures used in this guide. French-speaking readers should check quelllm.fr for EU-pricing equivalents.
Frequently Asked Questions
Is the RTX 5090 worth it over a used RTX 3090 in 2026?
Only if you need 32 GB of VRAM or you can't tolerate used-market risk. The 5090 is ~2.3× faster on 32B models but costs 3.5× more. For purely 30B-class workloads, two used 3090s deliver more VRAM and more aggregate bandwidth for less money.
Can I run Llama 3.3 70B on a single 24 GB GPU?
Only at Q2_K (~24 GB) with no context room, and quality degrades noticeably. Q3_K_M is ~30 GB and needs a 32 GB card (RTX 5090) or offloading to CPU/RAM (slow, 4–7 tok/s). For 70B at usable quality, plan for 48 GB+ of VRAM.
Why does Apple Silicon feel slow on long prompts even with high RAM?
Prefill (processing the input prompt) is compute-bound, not bandwidth-bound. Apple's GPU compute throughput is much lower than NVIDIA's. M4 Max processes a 16K prompt in ~9 seconds; an RTX 5090 does it in 1.5. For decode (output tokens), Apple is competitive because that phase is bandwidth-bound.
Is dual-GPU inference still worth setting up in 2026?
Yes, for 70B models on a budget. Two used 3090s ($1,400 total) beat a single 5090 ($2,650) on VRAM and equal it on aggregate bandwidth. Use vLLM tensor-parallel or llama.cpp with --split-mode row. The cost is power (700 W combined), noise, and a motherboard with two x8 PCIe slots.
Should I wait for the RTX 50-series Super refresh or Blackwell consumer Ti?
The RTX 5080 Super 24 GB launched March 2026 and is already in our table. No further consumer refresh is rumored before Q4 2026. If you need a GPU now, buy now — the next meaningful jump is the RTX 60-series in 2027, and waiting eight months for a hypothetical $200 discount is rarely rational.
Does VRAM bandwidth or capacity matter more?
Capacity is binary — either the model fits or it doesn't. Bandwidth is linear — more bandwidth means more tok/s. Always solve capacity first (pick a card that fits your target model), then optimize for bandwidth within that VRAM tier. There is no tok/s advantage if the model can't load.
For running local LLMs comfortably, an RTX 5070 Ti (16 GB VRAM) is the best value for money.
Amazon Check RTX 5070 Ti price →As an Amazon Associate, BestLLMfor earns from qualifying purchases, at no extra cost to you. It does not influence our independent rankings.