Beginner 12 minGPU

Which GPU for a local LLM? RTX 4070 to 5090, Mac (2026)

Direct response

For a local LLM, the #1 criterion is VRAM (not raw gaming performance), followed by tokens per second, then power consumption. A used RTX 3090 (24 GB VRAM) can outperform a new RTX 4070 for local AI because of its larger memory, despite its older architecture.

The GPU is the component that most changes your local AI experience. But you are not looking for the same GPU as a gamer: for an LLM, what matters is VRAM, then tokens per second, then power consumption. This guide untangles the 2026 offerings (NVIDIA, AMD, Apple Silicon, Intel Arc), gives real configurations by budget, and explains why a used RTX 3090 can beat a new 4070.

Choosing a machine? Our picks by budget →

By Mohamed Meguedmi·Update 2026-08-27·Tested on Windows, macOS, and Linux
Recommended hardware

Buying alternative for this guide: GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395).

A mini PC is a complete machine: check the available memory and engine compatibility. It does not replace macOS/MLX or CUDA.

Why this choice? Our complete guide on GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395) →

Compare all options by budget, from €800 to €3,500 →

Small budget: RTX 5060 · Large models: RTX 5090 · Mac Studio.

On the go: which laptop for local AI →

Affiliate links — commission possible at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.

#VRAM, the undisputed queen

An LLM must fit in VRAM to run at full speed. As soon as it spills into system RAM, speed collapses by a factor of 10. The size of the models you can run is therefore dictated by VRAM, NOT by the GPU’s raw power.

8 GB
Qwen 3.5 9B (6.6 GB). Comfortable usage, with long context limited by memory. Entry-level.
12 GB
Qwen 3.5 9B in Q8 (11 GB) with comfortable headroom, and a 24B in Q4 that fits. Sweet spot for price/performance.
16 GB
a comfortable 24B (Devstral, Mistral Small), gpt-oss 20B or a 27B in Q4 that fits. An excellent compromise.
24 GB
Qwen 3.8 27B (18 GB), which runs comfortably, the 30B-A3B code MoEs, and a 35B in Q4.
48 GB +
Comfortable 70B in Q4, LoRA fine-tuning possible, multiple models in parallel.
→
Golden rule
At the same budget, always choose the GPU with the most VRAM, even if its raw tokens/sec are slightly lower. A used 24 GB 3090 beats a new 12 GB 4070 for LLMs, despite the one-generation gap.

#Determine your needs

Occasional chat, simple tasks
Qwen 3.5 9B = ~7 GB of VRAM is enough. RTX 4060 Ti 16 GB, RTX 3060 12 GB, Arc A770 16 GB.
Dev (autocomplete + chat)
Qwen2.5-Coder 7B base for FIM autocompletion + Devstral 24B for tandem chat. 16–20 GB recommended.
Pro RAG, long documents
32k+ context consuming an additional 4–6 GB of VRAM. Aim for at least 16 GB.
Fine-tuning, experimentation
24 GB (RTX 3090/4090) minimum. For a true full fine-tune of a 7B, target 48 GB (A6000, RTX 6000 Ada).
Serve a team of 10 people
RTX 4090 or A6000 solo, or 2× RTX 4090 in tensor parallel for a 70B.

#NVIDIA — the 2026 options

RTX 3060 12 GB (~€250)
The budget option. CUDA, 12 GB VRAM, decent performance for 7B and 13B in Q4. Recommended entry-level choice.
RTX 4060 Ti 16 GB (~€450)
Perfect for LLMs: 16 GB of VRAM, low power consumption (160 W). Less powerful than a 3090 in raw performance, but more efficient.
RTX 4070 Super 12 GB (~€600)
Excellent at compute, but the 12 GB is limiting. Good if you also game.
RTX 3090 24 GB (~€700 used)
⭐ The best VRAM/€ value on the market. 2 years old, but the 24 GB work miracles. High power consumption (350 W).
RTX 4090 24 GB (used, variable price)
The best consumer GPU of 2023–2025. Performance + 24 GB. Stock is scarce in 2026; priority RTX 5090.
RTX 5080 16 GB / 5090 32 GB (≈ €1,500 to €1,850 / ≈ €5,700 in late September 2026)
2025 generation. 5090 = 32 GB VRAM, the dream of LLM users.
RTX A6000 48 GB (professional, ~€4,500)
48 GB of VRAM for 70B models in Q6, or fine-tuning a 13B model. For professionals who can amortize the cost.

#AMD — ROCm catch-up

RX 7600 XT 16 GB (~€350)
Entry-level AMD. 16 GB of VRAM for the price of a 3060. Official ROCm.
RX 7800 XT 16 GB (~€550)
4070-class performance, stable ROCm, excellent performance-per-dollar.
RX 7900 XTX 24 GB (~€900)
24 GB for much less than a used 3090. For those willing to accept ROCm's complexity.
RX 9070 XT 16 GB (≈ €1,070 at the end of September 2026)
New RDNA 4 generation (2025). Still needs evaluation on the ROCm side.
!
ROCm: still less seamless than CUDA
In 2026, ROCm works for LLM inference but remains less universal than CUDA. Some edge cases (exotic quantizations, newer backends) do not work. Patience required.

#Apple Silicon — a special case

M-series Macs don’t have a dedicated GPU but do have unified memory: the total RAM is available to the GPU. A 64 GB Mac can allocate 48 GB to an LLM—enough to run a 70B.

Mac mini M6 16 GB (1 049 €)
Usable minimum (the Mac mini M4 is no longer sold new). A 7B runs well. Ridiculously low power consumption (20 W during inference).
Mac mini M5 Pro 24 GB (€1,999, 48–64 GB in Apple configuration)
Excellent value for LLMs (replaces the M4 Pro). With 48 GB, ~36 GB allocated to the GPU: a 32B Q6 fits comfortably.
MacBook Pro M4 Max 64 GB (~4000 €)
Laptop + 70B Q4 LLM. Ergonomic equivalent of a RTX 4090 desktop + mobility.
Mac Studio M5 Ultra 96 GB (starting at €6,599, Apple price in late September 2026)
The top choice. About 72 GB usable by LLMs out of 96 GB, and more on versions with additional memory. It can run the largest open-weight models available today in Q8, or several Qwen 3.8 27B in parallel.
→
Mac = complete LLM workstation
A Mac Studio Ultra draws less power than a gaming PC and makes no noise. For an office or home office, it is a coherent package—whereas a 4090 PC build requires a case, a 1000W power supply, and cooling.

#Intel Arc — scrappy outsider

Arc A770 16 GB (~300 €)
16 GB of VRAM for €300. Raw performance is below equivalent NVIDIA/AMD cards, but the VRAM/€ ratio is unbeatable.
Arc B580 12 GB (≈ €430 in late September 2026)
New Battlemage generation. Much better performance. Promising.
i
Via Vulkan
llama.cpp compiled with Vulkan runs very well on Arc. oneAPI is also an option. Less documented than NVIDIA, but the community is growing.

#Concrete budgets

€300 — just to get started
RTX 3060 12 GB used, or Arc A770 16 GB. You can run 7B–13B without any problems.
€700—the sweet spot
RTX 3090 24 GB used. The best VRAM/€ ratio. Runs any 32B model and 70B models in Q4.
1500 € — professional comfort
A new RTX 5070 Ti 16 GB (≈1,250 to 1,400 €) to add to a PC that already has 64 GB of RAM.
€3000 — LLM workstation
A full PC with an RTX 5080 16 GB (≈ €1,500 to €1,850 for the card) OR a Mac Studio M5 Max 36 GB (€2,999). To load a 70B: a Ryzen AI Max+ 395 mini PC with 128 GB (≈ €3,300, up to €4,000 depending on the model). Makes sense for a developer who does this every day.
€6,600 and up—the very high end
Mac Studio M5 Ultra (96 GB and up, starting at €6,599) OR PC with 2× RTX 5090 (≈ €11,400 for both cards, late September 2026). LLM server for a small team.

#Buy new or used?

Nine
Warranty, latest generation, long-term driver support. For a professional investment that can be amortized over 3–5 years.
Use case (RTX 3090, 4090)
30–40% gain. LLM-oriented cards haven’t suffered like mining cards (most spent few hours under load).
Occasional checks
24-hour inference burn-in, temperature monitoring (must not exceed 80°C), no thermal pads replaced, original receipt preferred.
Business purchasing
New RTX A6000 or L40 cards. 3-year warranty, professional pricing, performance identical to consumer models. For those who need to amortize the cost.

Prices change quickly: every Monday and Thursday, our tracker records the lowest price for local AI graphics cards, along with the price per GB of VRAM.

Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.

Prices in euros (€) are French market prices including VAT, as checked by BestLLMfor. US prices differ: the Amazon buttons show the current US price.