Budget guide · updated 2026
Build your local LLM stack by budget
The cheapest hardware + model combination that actually runs Qwen 3, Gemma 3, and DeepSeek R1 distillations at a usable speed in 2026.
Which machine fits my budget?
Enter your budget and priority: the tool suggests machines that fit within them, what they can run, and their prices verified this week.
How to read this guide
Set two variables: budget et primary use. Everything else—quantization, runtime, context length—follows from it. Start with VRAM, choose the largest model that fits in Q4_K_M with headroom at 8K context, then choose the runtime.
Rule of thumb: a GGUF Q4_K_M model takes up approximately (paramètres × 0,6) Go of VRAM, plus ~1.5 GB of KV cache at 8K. A 14B therefore needs ~10 GB, which explains why the used RTX 3060 12 GB remains a price-to-performance benchmark.
Choose your tier
| Tier | Budget | Reference configuration | VRAM / unified memory | Ideal model (Q4_K_M) |
|---|---|---|---|---|
| Entry | 400-550 € | Used Ryzen 7 5700U mini PC, 16 GB DDR4 | 0 GB GPU / 16 GB RAM | Qwen3 4B |
| Value | 850-1 000 € | Ryzen 5 7600 + RTX 3060 12 GB used + 32 GB DDR5 | 12 GB | Qwen 3 14B or Qwen 2.5 Coder 14B |
| Performance | 1 400-1 600 € | Ryzen 7 7700 + RTX 3090 24 GB used + 64 GB DDR5 | 24 GB | Qwen3-Coder 30B-A3B or DeepSeek-R1-Distill-Qwen-32B |
| Silent | ≈ 2 000 € | Apple Mac mini M5 Pro 24 GB (new) | ≈ 16 GB usable | Qwen 3 14B |
Indicative used-market prices in France/EU as of mid-2026—the used market fluctuates; check current prices before buying.
Tier 2 — the €900 champion
This is the tier we recommend to most readers. A used RTX 3060 12 GB typically sells for around €200–240 on the secondary market, and 12 GB of VRAM is exactly the sweet spot for 14B-class models.
| Model (Q4_K_M) | VRAM used | Tokens/s (estimated) | Ideal for |
|---|---|---|---|
| Qwen 3 8B | ~6 GB | ~50 | Chat, RAG |
| Qwen 2.5 Coder 14B | ~10 GB | ~28 | Code completion, refactoring |
| Gemma 3 12B | ~8 GB | ~34 | Multilingual, document QA |
| DeepSeek-R1-Distill-Qwen-14B | ~10 GB | ~28 | Step-by-step reasoning |
Estimate, not a QuelLLM measurement: approximately 70% of the theoretical ceiling (the RTX 3060's bandwidth, 360 GB/s, divided by the model size). The ceiling for an 8B in Q4 is approximately 72 tokens/s, and for a 14B approximately 40.
Runtime: Ollama, llama.cpp, or vLLM?
- Ollama — the best choice for single-user use on a workstation. Automatically downloads GGUFs, exposes an OpenAI-compatible API, and integrates with Open WebUI in one command.
- llama.cpp —the best choice when you need to extract every token/s from a CPU or unified memory. More technical CLI, better portability.
- vLLM — best for serving multiple users or agents in parallel, thanks to PagedAttention and continuous batching.
Solo developers → Ollama by default. Teams of 3+ → vLLM behind a small FastAPI proxy.
Power consumption, noise, and total cost
| Config | Idle (W) | Load (W) | kWh/year (4 h/day under load) | Cost/year @ €0.20/kWh |
|---|---|---|---|---|
| Entry-level mini PC | 8 | 35 | 51 | 10,20 € |
| 3060 value | 45 | 220 | 321 | 64,20 € |
| 3090 performance | 50 | 360 | 526 | 105,10 € |
| Mac mini (Pro chip) | 6 | 65 | 95 | 19,00 € |
Power figures as orders of magnitude. Calculation: load power × 4 h × 365 days, with the machine turned off the rest of the time. If left idling, add idle power × 20 h × 365.
Even the most power-hungry configuration consumes about €105 in electricity per year at this rate—just over five months of a €20/month subscription.
Verdict
| Profile | Recommended tier | Why |
|---|---|---|
| Curious, chat only | €400 input | Qwen3 4B on CPU is faster than reading |
| Solo developer, daily coding | Value €900 | A 14B coding model at ~28 tok/s is the price/performance sweet spot |
| Reasoning, agents, multi-model | €1,500 performance | 24 GB unlocks 32B-class models |
| Silence and zero maintenance | Quiet ≈ €2,000 | The Mac mini consumes only a few watts at idle |
Frequently asked questions
Is 8 GB of VRAM enough for useful workloads in 2026?
Marginally. You can run Qwen3 8B Q4_K_M with a 4K context on an RTX 3050 8 GB or an RX 6600, but you lose the headroom for 14B models that make local inference truly worthwhile. Aim for 12 GB if possible.
Should you buy an AMD card instead of NVIDIA?
For inference only, Radeon RDNA3 and RDNA4 cards (7800 XT, 7900 GRE, RX 9070 XT) work well with llama.cpp's Vulkan backend and with ROCm. They often cost less for the same memory capacity, but you lose part of the CUDA ecosystem: more limited vLLM support, flash-attn, and most fine-tuning tools.
Can you fine-tune with these budgets?
LoRA fine-tuning of 7B models works on a 12 GB card with Unsloth. A 14B+ realistically requires 24 GB. Full fine-tuning beyond 1B exceeds what consumer hardware can support in 2026.
How does unified Mac memory compare with dedicated VRAM?
Its bandwidth remains well below that of a high-end card: 307 GB/s on an M5 Pro Mac mini versus 936 GB/s on a RTX 3090. For a model that fits on the card, the Mac therefore generates about 2 to 3 times fewer tokens per second. It wins on power consumption, memory capacity, and loading multiple models; it loses on tokens per second per euro.