Home›AI hardware›Which PC for local AI

Which PC for local AI? Our picks by budget

To run an AI model locally (Llama, Qwen, Mistral, gpt-oss…), a machine is chosen first according to its memory, which determines the model size, then according to its bandwidth, which determines their speed. Here are our picks for each budget, with the models that actually run on them.

Updated 09/10/2026 · indicative prices observed from our merchants in September 2026, which change quickly

By Mohamed Meguedmi · model sizes sourced from our catalog · our method

In brief: under €1,000, add a 16 GB card (RTX 5060 Ti) to your PC; at around €1,300, the RTX 5070 Ti twice as fast for the same memory; at around €2,000, a complete machine: Mac mini M5 Pro or Ryzen AI Max+ 395 64 GB mini-PC ; around €3,000–€3,500, 96 to 128 GB of unified memory for 70B models. Beyond that, a workstation such as the DGX Spark (273 GB/s, ≈ €6,100 to €6,600) generates more slowly than a much cheaper Mac Studio M5 Max.

Our picks at a glance

BudgetOur pickMemory for modelsUp to (Q4)Where to buy
Less than €1,000RTX 5060 Ti 16 GB
Graphics card · ≈ €820
16 GB14B handles it comfortably
1 200 – 1 500 €RTX 5070 Ti 16 GB
Graphics card · ≈ €1,300–€1,400
16 GB, 2× the bandwidth14B, faster
2 000 – 2 500 €Mac mini M5 Pro 24 GB / 512 GB
Complete machine · ≈ €2,000
≈ 16 GB (out of 24)14B
2 000 – 2 500 €GMKtec EVO-X2 mini PC 64 GB (Ryzen AI Max+ 395)
Complete machine
64 GB30B with room to spare

GMKtec: official European store (French-language site), with delivery in France averaging 5 to 8 days.

3 000 – 3 500 €Mac Studio M5 Max 96 GB (32-core GPU)
Workstation · ≈ €3,000
≈ 72 GB (out of 96)70B
3 300 – 4 000 €Minisforum MS-S1 Max mini PC, 128 GB / 2 TB (Ryzen AI Max+ 395)
Complete machine
128 GBgpt-oss 120B
3 600 – 4 600 €Comete IA prebuilt PC (RTX 5070 Ti 16 GB)
Assembled tower · ≈ €3,650
16 GB (+ 32 GB of RAM)14B handles it comfortably
3 600 – 4 600 €Meteora AI-built PC (RTX 5080, 64 GB of RAM)
Prebuilt tower
16 GB (+ 64 GB RAM)14B; 30B MoE partly in RAM
PortableRTX 5080 16 GB laptop
Laptop · starting at ≈ €2,500
16 GB14B handles it comfortably
PortableMacBook Pro M5 Max 36 GB
Laptop · starting at €4,199 (Apple price)
≈ 27 GB (out of 36)30B-class

Affiliate links — BestLLMfor may earn a commission from purchases at no extra cost to you, which does not influence these independently determined recommendations. As an Amazon Associate, BestLLMfor earns from qualifying purchases.

First: what model size do you want to run?

The machine's memory must hold the entire model, plus 1 to 2 GB for the context. The table shows the memory required in Q4, the most widely used quantization. On Mac, expect about 75% of unified memory is usable by the GPU, somewhat less so on base configurations (≈ 16 GB out of 24).
SizeExamplesQ4 memoryMinimum machine
7-8BLlama 3.1 8B≈ 6 GB8 GB card, or processor (slow)
12-14BGemma 3 12B, Qwen 3 14B7–9 GB12-16 GB card, 24 GB Mac
20-24Bgpt-oss 20B, Mistral Small 3.2 24B13–14 GB16 GB card (tight fit)
27-32BGemma 3 27B, Qwen 3 32B, Qwen 3 30B-A3B16–19 GB24–32 GB card, 36 GB and more unified memory
70-72BLlama 3.3 70B Instruct, Qwen 2.5 72B Instruct40-42 GB64 to 128 GB unified memory
120B (MoE)gpt-oss 120B≈ 70 GB96–128 GB unified memory

A model MoE (Mixture of Experts) activates only part of its parameters for each word: it is much faster than a dense model of the same size, but requires just as much memory to load. Do you already know which model you want? The configurator checks whether it fits on a machine.

The catch: more memory doesn't mean more speed

For a dense model, generation speed depends mainly on the memory bandwidth : for every word, the machine rereads the entire model. So a 16 GB graphics card is faster than a 128 GB workstation… on models that fit within its 16 GB.
MachineBandwidth
RTX 5080 (16 GB)960 GB/s
RTX 5070 Ti (16 GB)896 GB/s
Radeon RX 9070 XT (16 GB)644 GB/s
Apple M5 Max (36-128 GB)460 to 614 GB/s
RTX 5060 Ti (16 GB)448 GB/s
Mac mini M5 Pro (24-64 GB)307 GB/s
NVIDIA DGX Spark (128 GB)273 GB/s
Ryzen AI Max+ 395 (64-128 GB)≈ 212 GB/s

Real-world measurement: on a MacBook Pro M5 Max 128 GB, gpt-oss 120B (MoE) outputs 79.1 tokens/s, faster than a dense 27B model (36.6 tokens/s in GGUF), because it activates only a small portion of its parameters. View the full test.

Under €1,000: a 16 GB card for your PC

Already have a desktop PC? Add a RTX 5060 Ti 16 GB is the cheapest way to run models up to 14 billion parameters comfortably, and a 24B with tight quantization. Avoid its 8 GB version: 14B models no longer fit.

RTX 5060 Ti 16 GB

Graphics card to add to an existing PC

  • Memory16 GB of Dedicated GDDR7
  • Bandwidth448 GB/s
  • Indicative price≈ 820 €
  • Up to14B models at ease, 24B with tight quantization

What runs on it (memory required in Q4)

Reference pointA dedicated card remains the fastest solution per euro at the same model size. No QuelLLM measurements have been published for this card yet.

✓ The cheapest way to exceed 8 GB of VRAM · CUDA: compatible with Ollama, LM Studio, llama.cpp, and vLLM

✗ 16 GB ceiling: no comfortable 32B · Assumes a recent PC with a compatible power supply and case

AlternativeRadeon RX 9070 XT 16 GB (≈ €960–€1,070) — Still 16 GB, with higher bandwidth (644 GB/s); on the software side, ROCm remains less mature than CUDA: check your tools. Our fact sheet →

No graphics card at all? A small 3 to 8B model can also run on the processor, slowly: see our GPU-free LLM guide.

€1,200 to €1,500: the fastest 16 GB card

La RTX 5070 Ti keeps 16 GB of memory but doubles the bandwidth (896 GB/s versus 448 GB/s): the same models, generated significantly faster. This is the right choice for everyday coding or RAG.

RTX 5070 Ti 16 GB

Graphics card to add to an existing PC

  • Memory16 GB of Dedicated GDDR7
  • Bandwidth896 GB/s (twice the 5060 Ti)
  • Indicative price≈ 1 300 – 1 400 €
  • Up toSame models as the 5060 Ti (16 GB), generated noticeably faster

What runs on it (memory required in Q4)

Reference pointA dense model’s generation speed follows memory bandwidth: at 896 GB/s, the 5070 Ti doubles the 5060 Ti’s bandwidth (448 GB/s) with the same memory.

✓ The best speed-to-price ratio at 16 GB · CUDA: all tools and LoRA fine-tuning work

✗ Still 16 GB: same size ceiling as the 5060 Ti · Card-only price, excluding the rest of the PC

AlternativeRTX 5080 16 GB (≈ €1,700–€1,850) — 960 GB/s, only about 7% more bandwidth than the 5070 Ti, at a significantly higher price: memory remains at 16 GB.

€2,000 to €2,500: a complete, quiet machine

Don't need to upgrade your PC? Two complete machines share this budget: the Mac mini M5 Pro 24 GB for macOS and silence, or a Ryzen AI Max+ 395 mini PC with 64 GB for 2.7 times more memory and 30B models.

Mac mini M5 Pro 24 GB / 512 GB

Complete machine, macOS

  • Memory24 GB unified, of which ≈ 16 GB can be used by the GPU
  • Bandwidth307 GB/s
  • Indicative price≈ 2 000 €
  • Up to14B models that are comfortable

What runs on it (memory required in Q4)

Reference pointOn Mac, MLX often accelerates generation: +86% measured on a 27B model (M5 Max, a reader's test using our protocol).

✓ Compact, quiet, low power consumption · MLX, Ollama, and LM Studio optimized for Apple Silicon

✗ 24 GB quickly becomes limiting beyond 14B · No CUDA; memory cannot be upgraded after purchase

GMKtec EVO-X2 mini PC 64 GB (Ryzen AI Max+ 395)

Complete machine, Windows or Linux

  • Memory64 GB unified memory shared by the CPU and GPU
  • Bandwidth≈ 212 GB/s
  • Up to30B class runs comfortably; a 70B in Q4 fits but remains slow

What runs on it (memory required in Q4)

EstimateVery fast on MoE models (Qwen 3 30B-A3B: estimated 50 to 100 tokens/s); a dense 70B drops to 3-5 tokens/s (estimates from our profile, not a measurement).

✓ 2.7 times the Mac mini's memory for a similar price · Ideal for 30B MoE models

✗ No CUDA or macOS: AMD tools need verification · Modest bandwidth for large dense models

GMKtec: official European store (French-language site), with delivery in France averaging 5 to 8 days.

€3,000 to €3,500: 96 to 128 GB for 70B models

For a 70B model (Llama 3.3, Qwen 2.5 72B) or a large MoE such as gpt-oss 120B, you need 64 to 128 GB of unified memory. The Mac Studio M5 Max 96 GB generates faster; a Ryzen AI Max+ 395 mini-PC with 128 GB handles larger models: the Minisforum MS-S1 Max on Amazon, or the GMKtec EVO-X2 128 GB on the official GMKtec store.

Mac Studio M5 Max 96 GB (32-core GPU)

Compact workstation, macOS · configuration sold by Materiel.net (ref. MHL64ZD/A)

  • Memory96 GB unified, of which ≈ 72 GB is available to the GPU
  • Bandwidth460 GB/s
  • Indicative price≈ 3 000 €
  • Up to70B models in Q4; gpt-oss 120B (70 GB) too tight

What runs on it (memory required in Q4)

Comparable metricOn a MacBook Pro M5 Max 128 GB (40 cores, 614 GB/s): Qwen 3.8 27B at 36.6 tokens/s in GGUF and 68 in MLX. Expect slightly less with 32 cores (460 GB/s).

✓ The fastest machines with 96–128 GB in this price range · Quiet, with the MLX ecosystem

✗ No CUDA · Non-expandable memory: choose the right configuration when you buy

Minisforum MS-S1 Max mini PC, 128 GB / 2 TB (Ryzen AI Max+ 395)

Complete machine, Windows or Linux · two 10 GbE ports, PCIe x16 slot

  • Memory128 GB unified and shared by CPU / GPU
  • Bandwidth≈ 212 GB/s
  • Up toUp to gpt-oss 120B (MoE); 70B loadable in Q4

What runs on it (memory required in Q4)

EstimateA dense 70B runs at 3-5 tokens/s, while MoE models are much faster (estimates from our Ryzen AI Max spec sheet, not a measurement).

✓ One-third more memory than the 96 GB Mac Studio · Loads the largest mainstream models (gpt-oss 120B)

✗ Half the bandwidth of the Mac Studio · No CUDA or macOS

€3,600 to €4,600: a complete, assembled, and tested tower

Would you prefer a ready-to-use desktop tower over a card you have to install? Materiel.net assembles PCs dedicated to AI: the Comete IA (RTX 5070 Ti) and the Meteora IA (RTX 5080, 64 GB of RAM). They keep 16 GB of VRAM : the same model sizes as a single card, without the assembly.

Comete IA prebuilt PC (RTX 5070 Ti 16 GB)

Built and tested by Materiel.net · Ryzen 7 9800X3D, 2 TB SSD, Windows 11

  • Memory16 GB of VRAM + 32 GB of DDR5 RAM
  • Bandwidth896 GB/s (graphics card)
  • Indicative price≈ 3 650 €
  • Up to14B models run comfortably, as does the card alone

What runs on it (memory required in Q4)

Reference pointSame card as the €1,200–1,500 tier: you are mainly paying for a complete, assembled, tested machine.

✓ Ready to use, no assembly required · High-end processor and SSD for the rest of your workloads

✗ Still 16 GB of VRAM · Much more expensive than a single card if you already have a PC

Meteora AI-built PC (RTX 5080, 64 GB of RAM)

Built and tested by Materiel.net · Ryzen 7 9800X3D, 2 TB SSD, Windows 11

  • Memory16 GB of VRAM + 64 GB of DDR5 RAM
  • Bandwidth960 GB/s (graphics card)
  • Up to14B runs comfortably on the card; beyond that, part of the model spills into RAM, much more slowly

What runs on it (memory required in Q4)

Reference pointThe 64 GB of RAM makes it possible to load an MoE such as Qwen 3 30B-A3B partially outside the GPU, at a significantly reduced speed.

✓ Fastest 16 GB card + 64 GB RAM · Ready to use, no assembly required

✗ 16 GB of VRAM: no fast 32B · At the same price, a mini PC or a Mac with unified memory loads larger models

On the go: RTX 5080 laptop or MacBook Pro M5 Max?

Un RTX 5080 16 GB laptop retains CUDA and the speed of a dedicated card, within the 16 GB limit. The MacBook Pro M5 Max offers more memory and quiet operation, but the real leap (70B, 120B) requires its 64–128 GB configurations, which are much more expensive.

RTX 5080 16 GB laptop

Windows laptop

  • Memory16 GB of Dedicated GDDR7
  • Bandwidth≈ 700 to 900 GB/s depending on the chassis
  • Indicative pricestarting at ≈ €2,500
  • Up to14B models run comfortably, 24B is tight

What runs on it (memory required in Q4)

Reference pointSame principle as a 16 GB desktop card, slightly throttled by the laptop's thermal envelope.

✓ CUDA on the go · The cheapest AI laptop with 16 GB of VRAM

✗ Noise and heat during long generations · 16 GB ceiling

MacBook Pro M5 Max 36 GB

Portable macOS

  • Memory36 GB unified, of which ≈ 27 GB can be used by the GPU
  • Bandwidth460 GB/s
  • Indicative pricestarting at €4,199 (Apple price)
  • Up to30B class in Q4

What runs on it (memory required in Q4)

MetricIn the 128 GB version (40 cores, desktop): gpt-oss 120B at 79,1 tokens/s and Qwen 3.8 27B at 68 tokens/s in MLX (reader test using our protocol, 16/09/2026).

✓ The only laptop that supports the M5 Max chip · Quiet, battery life

✗ Expensive for 36 GB: 64–128 GB requires custom configuration Apple · No CUDA

DGX Spark, RTX 5090, Mac Studio M5 Ultra: should you aim higher?

Rarely for personal use. Each of these machines has a specific advantage, but none offers the best price-to-generation-speed ratio today.
MachineMemory · bandwidthIndicative priceOur take
NVIDIA DGX Spark128 GB · 273 GB/s≈ 6 100 – 6 600 €For the NVIDIA ecosystem and training. In generation, its bandwidth (273 GB/s) remains below that of a Mac Studio M5 Max and barely above that of a 128 GB mini-PC costing half as much. Same GB10 chip, delivered faster: the MSI EdgeXpert (128 GB).
RTX 509032 GB GDDR7highThe fastest consumer card: 32 GB to load a 30B model in full. For big budgets, in a PC capable of powering it.
Mac Studio M5 Ultra96 to 512 GB · up to 1.2 TB/sfrom €6,599For very large models (120B and above, giant MoE models) at good speed.

Where to find themMSI EdgeXpert (DGX Spark's GB10 chip, 128 GB)

RTX 5090 32 GB (card only)

Mac, PC NVIDIA, or AMD mini PC: what changes on the software side

PlatformToolsGood to know
PC with a NVIDIA card (CUDA)Ollama, LM Studio, llama.cpp, vLLMEverything works, including training (LoRA). The simplest choice.
Mac Apple SiliconOllama, LM Studio, llama.cpp, MLXMLX often accelerates generation (+86% measured on a 27B). No CUDA.
Ryzen AI Max mini PC, Radeon cardOllama, LM Studio, llama.cpp (Vulkan or ROCm)Compatibility must be verified depending on the tool; ROCm remains less mature than CUDA.

Frequently asked questions

What is the best PC for local AI in 2026?

It depends on the size of the models you’re targeting. Up to 14B, a 16 GB graphics card (RTX 5060 Ti, or RTX 5070 Ti for speed) in an existing PC is sufficient. For a complete machine, use a Mac mini M5 Pro (14B) or a Ryzen AI Max+ 395 64 GB mini PC (30B). For 70B models, you need 64 to 128 GB of unified memory: a Mac Studio M5 Max or a 128 GB mini PC.

Which graphics card should you choose for local AI?

Target at least 16 GB of VRAM: the RTX 5060 Ti 16 GB for the budget, the RTX 5070 Ti for speed (896 GB/s), the Radeon RX 9070 XT if AMD tools work for you. Avoid the 8 GB versions. Details in our guide to choosing your GPU.

Which mini PC should you choose for AI?

A unified-memory mini-PC: Ryzen AI Max+ 395 with 64 GB (GMKtec EVO-X2) or 128 GB (GMKtec EVO-X2 128 GB, Minisforum MS-S1 Max), or a Mac mini M5 Pro if you want macOS. A standard mini-PC without substantial shared memory can run only small models.

Do you need a graphics card or a unified-memory mini PC?

The graphics card is faster (448 to 960 GB/s) but limited to 16 GB in consumer hardware. Unified memory (64 to 128 GB) loads much larger models, but generates more slowly (212 to 614 GB/s depending on the machine). Choose based on the target model size.

How much does a PC for running an LLM locally cost?

About €750 to €850 for a 16 GB card added to an existing PC, €2,000 for a complete machine that runs 14B to 30B models, and €3,000 to €3,500 for 70B and larger models.

Can you run an LLM without a graphics card?

Yes, small 3B to 8B models run on the processor with 16 GB of RAM, but slowly. See our GPU-free LLM guide.

Mac or PC for local AI?

The Mac offers lots of unified memory, silence, and MLX (+86% speed measured on a 27B), but no CUDA. A NVIDIA PC is faster per dollar and compatible with everything, including training. An AMD mini PC offers the most memory for the price, with tools that need to be verified.

What are the drawbacks of a mini PC for AI?

Modest bandwidth—about 212 GB/s on the Ryzen AI Max+ 395—means a dense 70B model runs at 3–5 tokens/s by our estimates. No CUDA, and memory cannot be expanded after purchase.

Go further

Prices in euros (€) are French market prices including VAT, checked by QuelLLM. US prices differ: the Amazon buttons show the current US price.