Which PC for local AI? Our picks by budget
To run an AI model locally (Llama, Qwen, Mistral, gpt-oss…), a machine is chosen first according to its memory, which determines the model size, then according to its bandwidth, which determines their speed. Here are our picks for each budget, with the models that actually run on them.
Updated 09/10/2026 · indicative prices observed from our merchants in September 2026, which change quickly
Our picks at a glance
| Budget | Our pick | Memory for models | Up to (Q4) | Where to buy |
|---|---|---|---|---|
| Less than €1,000 | RTX 5060 Ti 16 GB Graphics card · ≈ €820 | 16 GB | 14B handles it comfortably | |
| 1 200 – 1 500 € | RTX 5070 Ti 16 GB Graphics card · ≈ €1,300–€1,400 | 16 GB, 2× the bandwidth | 14B, faster | |
| 2 000 – 2 500 € | Mac mini M5 Pro 24 GB / 512 GB Complete machine · ≈ €2,000 | ≈ 16 GB (out of 24) | 14B | |
| 2 000 – 2 500 € | GMKtec EVO-X2 mini PC 64 GB (Ryzen AI Max+ 395) Complete machine | 64 GB | 30B with room to spare | GMKtec: official European store (French-language site), with delivery in France averaging 5 to 8 days. |
| 3 000 – 3 500 € | Mac Studio M5 Max 96 GB (32-core GPU) Workstation · ≈ €3,000 | ≈ 72 GB (out of 96) | 70B | |
| 3 300 – 4 000 € | Minisforum MS-S1 Max mini PC, 128 GB / 2 TB (Ryzen AI Max+ 395) Complete machine | 128 GB | gpt-oss 120B | |
| 3 600 – 4 600 € | Comete IA prebuilt PC (RTX 5070 Ti 16 GB) Assembled tower · ≈ €3,650 | 16 GB (+ 32 GB of RAM) | 14B handles it comfortably | |
| 3 600 – 4 600 € | Meteora AI-built PC (RTX 5080, 64 GB of RAM) Prebuilt tower | 16 GB (+ 64 GB RAM) | 14B; 30B MoE partly in RAM | |
| Portable | RTX 5080 16 GB laptop Laptop · starting at ≈ €2,500 | 16 GB | 14B handles it comfortably | |
| Portable | MacBook Pro M5 Max 36 GB Laptop · starting at €4,199 (Apple price) | ≈ 27 GB (out of 36) | 30B-class |
Affiliate links — BestLLMfor may earn a commission from purchases at no extra cost to you, which does not influence these independently determined recommendations. As an Amazon Associate, BestLLMfor earns from qualifying purchases.
First: what model size do you want to run?
| Size | Examples | Q4 memory | Minimum machine |
|---|---|---|---|
| 7-8B | Llama 3.1 8B | ≈ 6 GB | 8 GB card, or processor (slow) |
| 12-14B | Gemma 3 12B, Qwen 3 14B | 7–9 GB | 12-16 GB card, 24 GB Mac |
| 20-24B | gpt-oss 20B, Mistral Small 3.2 24B | 13–14 GB | 16 GB card (tight fit) |
| 27-32B | Gemma 3 27B, Qwen 3 32B, Qwen 3 30B-A3B | 16–19 GB | 24–32 GB card, 36 GB and more unified memory |
| 70-72B | Llama 3.3 70B Instruct, Qwen 2.5 72B Instruct | 40-42 GB | 64 to 128 GB unified memory |
| 120B (MoE) | gpt-oss 120B | ≈ 70 GB | 96–128 GB unified memory |
A model MoE (Mixture of Experts) activates only part of its parameters for each word: it is much faster than a dense model of the same size, but requires just as much memory to load. Do you already know which model you want? The configurator checks whether it fits on a machine.
The catch: more memory doesn't mean more speed
| Machine | Bandwidth |
|---|---|
| RTX 5080 (16 GB) | 960 GB/s |
| RTX 5070 Ti (16 GB) | 896 GB/s |
| Radeon RX 9070 XT (16 GB) | 644 GB/s |
| Apple M5 Max (36-128 GB) | 460 to 614 GB/s |
| RTX 5060 Ti (16 GB) | 448 GB/s |
| Mac mini M5 Pro (24-64 GB) | 307 GB/s |
| NVIDIA DGX Spark (128 GB) | 273 GB/s |
| Ryzen AI Max+ 395 (64-128 GB) | ≈ 212 GB/s |
Real-world measurement: on a MacBook Pro M5 Max 128 GB, gpt-oss 120B (MoE) outputs 79.1 tokens/s, faster than a dense 27B model (36.6 tokens/s in GGUF), because it activates only a small portion of its parameters. View the full test.
Under €1,000: a 16 GB card for your PC
RTX 5060 Ti 16 GB
Graphics card to add to an existing PC
- Memory16 GB of Dedicated GDDR7
- Bandwidth448 GB/s
- Indicative price≈ 820 €
- Up to14B models at ease, 24B with tight quantization
What runs on it (memory required in Q4)
Reference pointA dedicated card remains the fastest solution per euro at the same model size. No QuelLLM measurements have been published for this card yet.
✓ The cheapest way to exceed 8 GB of VRAM · CUDA: compatible with Ollama, LM Studio, llama.cpp, and vLLM
✗ 16 GB ceiling: no comfortable 32B · Assumes a recent PC with a compatible power supply and case
Guide: which LLMs on a RTX 5060 Ti →Which LLM with 16 GB of VRAM →
AlternativeRadeon RX 9070 XT 16 GB (≈ €960–€1,070) — Still 16 GB, with higher bandwidth (644 GB/s); on the software side, ROCm remains less mature than CUDA: check your tools. Our fact sheet →
No graphics card at all? A small 3 to 8B model can also run on the processor, slowly: see our GPU-free LLM guide.
€1,200 to €1,500: the fastest 16 GB card
RTX 5070 Ti 16 GB
Graphics card to add to an existing PC
- Memory16 GB of Dedicated GDDR7
- Bandwidth896 GB/s (twice the 5060 Ti)
- Indicative price≈ 1 300 – 1 400 €
- Up toSame models as the 5060 Ti (16 GB), generated noticeably faster
What runs on it (memory required in Q4)
Reference pointA dense model’s generation speed follows memory bandwidth: at 896 GB/s, the 5070 Ti doubles the 5060 Ti’s bandwidth (448 GB/s) with the same memory.
✓ The best speed-to-price ratio at 16 GB · CUDA: all tools and LoRA fine-tuning work
✗ Still 16 GB: same size ceiling as the 5060 Ti · Card-only price, excluding the rest of the PC
Guide: which LLMs on a RTX 5070 Ti →Choose your GPU for local AI →
AlternativeRTX 5080 16 GB (≈ €1,700–€1,850) — 960 GB/s, only about 7% more bandwidth than the 5070 Ti, at a significantly higher price: memory remains at 16 GB.
€2,000 to €2,500: a complete, quiet machine
Mac mini M5 Pro 24 GB / 512 GB
Complete machine, macOS
- Memory24 GB unified, of which ≈ 16 GB can be used by the GPU
- Bandwidth307 GB/s
- Indicative price≈ 2 000 €
- Up to14B models that are comfortable
What runs on it (memory required in Q4)
Reference pointOn Mac, MLX often accelerates generation: +86% measured on a 27B model (M5 Max, a reader's test using our protocol).
✓ Compact, quiet, low power consumption · MLX, Ollama, and LM Studio optimized for Apple Silicon
✗ 24 GB quickly becomes limiting beyond 14B · No CUDA; memory cannot be upgraded after purchase
GMKtec EVO-X2 mini PC 64 GB (Ryzen AI Max+ 395)
Complete machine, Windows or Linux
- Memory64 GB unified memory shared by the CPU and GPU
- Bandwidth≈ 212 GB/s
- Up to30B class runs comfortably; a 70B in Q4 fits but remains slow
What runs on it (memory required in Q4)
EstimateVery fast on MoE models (Qwen 3 30B-A3B: estimated 50 to 100 tokens/s); a dense 70B drops to 3-5 tokens/s (estimates from our profile, not a measurement).
✓ 2.7 times the Mac mini's memory for a similar price · Ideal for 30B MoE models
✗ No CUDA or macOS: AMD tools need verification · Modest bandwidth for large dense models
GMKtec: official European store (French-language site), with delivery in France averaging 5 to 8 days.
€3,000 to €3,500: 96 to 128 GB for 70B models
Mac Studio M5 Max 96 GB (32-core GPU)
Compact workstation, macOS · configuration sold by Materiel.net (ref. MHL64ZD/A)
- Memory96 GB unified, of which ≈ 72 GB is available to the GPU
- Bandwidth460 GB/s
- Indicative price≈ 3 000 €
- Up to70B models in Q4; gpt-oss 120B (70 GB) too tight
What runs on it (memory required in Q4)
Comparable metricOn a MacBook Pro M5 Max 128 GB (40 cores, 614 GB/s): Qwen 3.8 27B at 36.6 tokens/s in GGUF and 68 in MLX. Expect slightly less with 32 cores (460 GB/s).
✓ The fastest machines with 96–128 GB in this price range · Quiet, with the MLX ecosystem
✗ No CUDA · Non-expandable memory: choose the right configuration when you buy
Minisforum MS-S1 Max mini PC, 128 GB / 2 TB (Ryzen AI Max+ 395)
Complete machine, Windows or Linux · two 10 GbE ports, PCIe x16 slot
- Memory128 GB unified and shared by CPU / GPU
- Bandwidth≈ 212 GB/s
- Up toUp to gpt-oss 120B (MoE); 70B loadable in Q4
What runs on it (memory required in Q4)
EstimateA dense 70B runs at 3-5 tokens/s, while MoE models are much faster (estimates from our Ryzen AI Max spec sheet, not a measurement).
✓ One-third more memory than the 96 GB Mac Studio · Loads the largest mainstream models (gpt-oss 120B)
✗ Half the bandwidth of the Mac Studio · No CUDA or macOS
€3,600 to €4,600: a complete, assembled, and tested tower
Comete IA prebuilt PC (RTX 5070 Ti 16 GB)
Built and tested by Materiel.net · Ryzen 7 9800X3D, 2 TB SSD, Windows 11
- Memory16 GB of VRAM + 32 GB of DDR5 RAM
- Bandwidth896 GB/s (graphics card)
- Indicative price≈ 3 650 €
- Up to14B models run comfortably, as does the card alone
What runs on it (memory required in Q4)
Reference pointSame card as the €1,200–1,500 tier: you are mainly paying for a complete, assembled, tested machine.
✓ Ready to use, no assembly required · High-end processor and SSD for the rest of your workloads
✗ Still 16 GB of VRAM · Much more expensive than a single card if you already have a PC
Meteora AI-built PC (RTX 5080, 64 GB of RAM)
Built and tested by Materiel.net · Ryzen 7 9800X3D, 2 TB SSD, Windows 11
- Memory16 GB of VRAM + 64 GB of DDR5 RAM
- Bandwidth960 GB/s (graphics card)
- Up to14B runs comfortably on the card; beyond that, part of the model spills into RAM, much more slowly
What runs on it (memory required in Q4)
Reference pointThe 64 GB of RAM makes it possible to load an MoE such as Qwen 3 30B-A3B partially outside the GPU, at a significantly reduced speed.
✓ Fastest 16 GB card + 64 GB RAM · Ready to use, no assembly required
✗ 16 GB of VRAM: no fast 32B · At the same price, a mini PC or a Mac with unified memory loads larger models
On the go: RTX 5080 laptop or MacBook Pro M5 Max?
RTX 5080 16 GB laptop
Windows laptop
- Memory16 GB of Dedicated GDDR7
- Bandwidth≈ 700 to 900 GB/s depending on the chassis
- Indicative pricestarting at ≈ €2,500
- Up to14B models run comfortably, 24B is tight
What runs on it (memory required in Q4)
Reference pointSame principle as a 16 GB desktop card, slightly throttled by the laptop's thermal envelope.
✓ CUDA on the go · The cheapest AI laptop with 16 GB of VRAM
✗ Noise and heat during long generations · 16 GB ceiling
MacBook Pro M5 Max 36 GB
Portable macOS
- Memory36 GB unified, of which ≈ 27 GB can be used by the GPU
- Bandwidth460 GB/s
- Indicative pricestarting at €4,199 (Apple price)
- Up to30B class in Q4
What runs on it (memory required in Q4)
MetricIn the 128 GB version (40 cores, desktop): gpt-oss 120B at 79,1 tokens/s and Qwen 3.8 27B at 68 tokens/s in MLX (reader test using our protocol, 16/09/2026).
✓ The only laptop that supports the M5 Max chip · Quiet, battery life
✗ Expensive for 36 GB: 64–128 GB requires custom configuration Apple · No CUDA
DGX Spark, RTX 5090, Mac Studio M5 Ultra: should you aim higher?
| Machine | Memory · bandwidth | Indicative price | Our take |
|---|---|---|---|
| NVIDIA DGX Spark | 128 GB · 273 GB/s | ≈ 6 100 – 6 600 € | For the NVIDIA ecosystem and training. In generation, its bandwidth (273 GB/s) remains below that of a Mac Studio M5 Max and barely above that of a 128 GB mini-PC costing half as much. Same GB10 chip, delivered faster: the MSI EdgeXpert (128 GB). |
| RTX 5090 | 32 GB GDDR7 | high | The fastest consumer card: 32 GB to load a 30B model in full. For big budgets, in a PC capable of powering it. |
| Mac Studio M5 Ultra | 96 to 512 GB · up to 1.2 TB/s | from €6,599 | For very large models (120B and above, giant MoE models) at good speed. |
Where to find themMSI EdgeXpert (DGX Spark's GB10 chip, 128 GB)
RTX 5090 32 GB (card only)
Mac, PC NVIDIA, or AMD mini PC: what changes on the software side
| Platform | Tools | Good to know |
|---|---|---|
| PC with a NVIDIA card (CUDA) | Ollama, LM Studio, llama.cpp, vLLM | Everything works, including training (LoRA). The simplest choice. |
| Mac Apple Silicon | Ollama, LM Studio, llama.cpp, MLX | MLX often accelerates generation (+86% measured on a 27B). No CUDA. |
| Ryzen AI Max mini PC, Radeon card | Ollama, LM Studio, llama.cpp (Vulkan or ROCm) | Compatibility must be verified depending on the tool; ROCm remains less mature than CUDA. |
Frequently asked questions
What is the best PC for local AI in 2026?
It depends on the size of the models you’re targeting. Up to 14B, a 16 GB graphics card (RTX 5060 Ti, or RTX 5070 Ti for speed) in an existing PC is sufficient. For a complete machine, use a Mac mini M5 Pro (14B) or a Ryzen AI Max+ 395 64 GB mini PC (30B). For 70B models, you need 64 to 128 GB of unified memory: a Mac Studio M5 Max or a 128 GB mini PC.
Which graphics card should you choose for local AI?
Target at least 16 GB of VRAM: the RTX 5060 Ti 16 GB for the budget, the RTX 5070 Ti for speed (896 GB/s), the Radeon RX 9070 XT if AMD tools work for you. Avoid the 8 GB versions. Details in our guide to choosing your GPU.
Which mini PC should you choose for AI?
A unified-memory mini-PC: Ryzen AI Max+ 395 with 64 GB (GMKtec EVO-X2) or 128 GB (GMKtec EVO-X2 128 GB, Minisforum MS-S1 Max), or a Mac mini M5 Pro if you want macOS. A standard mini-PC without substantial shared memory can run only small models.
Do you need a graphics card or a unified-memory mini PC?
The graphics card is faster (448 to 960 GB/s) but limited to 16 GB in consumer hardware. Unified memory (64 to 128 GB) loads much larger models, but generates more slowly (212 to 614 GB/s depending on the machine). Choose based on the target model size.
How much does a PC for running an LLM locally cost?
About €750 to €850 for a 16 GB card added to an existing PC, €2,000 for a complete machine that runs 14B to 30B models, and €3,000 to €3,500 for 70B and larger models.
Can you run an LLM without a graphics card?
Yes, small 3B to 8B models run on the processor with 16 GB of RAM, but slowly. See our GPU-free LLM guide.
Mac or PC for local AI?
The Mac offers lots of unified memory, silence, and MLX (+86% speed measured on a 27B), but no CUDA. A NVIDIA PC is faster per dollar and compatible with everything, including training. An AMD mini PC offers the most memory for the price, with tools that need to be verified.
What are the drawbacks of a mini PC for AI?
Modest bandwidth—about 212 GB/s on the Ryzen AI Max+ 395—means a dense 70B model runs at 3–5 tokens/s by our estimates. No CUDA, and memory cannot be expanded after purchase.
Go further
Prices in euros (€) are French market prices including VAT, checked by QuelLLM. US prices differ: the Amazon buttons show the current US price.