Which LLM on RTX 4090 (24 GB) ?
The RTX 4090 (October 2022) remains the benchmark for mainstream local AI in 2026: 24 GB of GDDR6X VRAM, 1008 GB/s of bandwidth, and 16,384 CUDA Ada Lovelace cores. It runs Qwen 3.5 9B at over 90 tok/s, a Qwen 3.8 27B Q4 comfortably, and even a Llama 70B Q4 with light offload. Now selling for €1700–2000 used (ex-LHR stock), it offers a better performance-per-euro ratio than a new 5080 for pure LLM use. This guide details all the models it can run and approximate throughput figures.
Choosing a machine? Our picks by budget →
Buying alternative for this guide: GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395).
A mini PC is a complete machine: check the available memory and engine compatibility. It does not replace macOS/MLX or CUDA.
Why this choice? Our complete guide on GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395) →
Compare all options by budget, from €800 to €3,500 →
On the go: which laptop for local AI →
Affiliate links — commission possible at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.
#The RTX 4090 in 2026, still the queen
- Architecture
- Ada Lovelace (AD102), manufactured on TSMC 4N. 76 billion transistors.
- VRAM
- 24 GB GDDR6X at 21 Gbps, 384-bit bus. 1008 GB/s bandwidth.
- CUDA cores
- 16,384. 4th-gen Tensor Cores with native FP8. 3rd-gen RT Cores.
- TDP
- 450 W. An 850 W PSU is recommended.
- 2026 price
- 1700–2000 € in good used condition.
#1. Compatible models 24 GB
| Model | Quant | VRAM | Tokens/sec | Usage |
|---|---|---|---|---|
| gpt-oss 20B | Q4_K_M | 13 GB | 117-143 | Instant chat |
| Devstral Small 2 24B | Q4_K_M | 14 GB | 36-44 | Chat + RAG FR |
| Qwen 3.6 27B | Q4_K_M | 16 GB | 29-35 | Pro code assistant |
| Qwen 3.8 27B | Q4_K_M | 16 GB | 20-24 | Quality reasoning |
| Mistral Small 24B | Q6_K | 20 GB | 32-38 | Versatile sweet spot |
| GLM 4.7 Flash | Q4_K_M | 19 GB | 90-110 | Advanced reasoning |
| Granite 4.1 30B Instruct | Q4_K_M | 17 GB | 27-33 | Complex multi-file code |
| Gemma 4 26B-A4B MoE | Q5_K_M | 19 GB | 54-66 | 2 GB RAM offload, usable |
| Gemma 4 31B | Q5_K_M | 22 GB | 27-33 | All in VRAM, degraded quality |
#2. Installation & CUDA
- 01NVIDIA 550+ driversCUDA 12.4+. For a new 4090, GeForce Experience handles it. Linux: nvidia-driver-565 or later.
- 02OllamaStandard installer from ollama.com. CUDA GPU auto-detection.
- 03First 24 GB model testollama run qwen3.8:27b — profite de la VRAM dispo et met en valeur la carte.
#3. Tokens/sec benchmarks
| Model | Q4_K_M | Q5_K_M | Q6_K | Q8_0 |
|---|---|---|---|---|
| Qwen 3.5 4B | 135 t/s | 122 t/s | 110 t/s | 90 t/s |
| Granite 4.2 8B | 100 t/s | 92 t/s | 82 t/s | 66 t/s |
| Qwen 3.5 9B | 92 t/s | 85 t/s | 76 t/s | 62 t/s |
| Gemma 4 12B | 70 t/s | 64 t/s | 57 t/s | 48 t/s |
| Mistral Small 24B | 42 t/s | 38 t/s | 34 t/s | 28 t/s |
| Qwen 3.8 27B | 32 t/s | 28 t/s | 25 t/s | — |
#4. Ada optimizations
- Flash Attention 2
- Enabled by default. FA3 is coming in future llama.cpp versions.
- KV cache Q8
- OLLAMA_KV_CACHE_TYPE=q8_0. Enables a 32k context on 14B in ~15 GB, leaving 9 GB for other processes.
- Power limit
- nvidia-smi -pl 350 (starting at 350 W). LLM: -2% speed, -25% heat and noise. Sustainable 24/7 configuration.
- FP8 via TensorRT-LLM
- Ada supports native FP8. Via TensorRT-LLM, +40% batch throughput on Qwen 3.5 9B. Especially useful for multi-user serving.
#5. 4090 vs 5090 vs 3090
| Criterion | RTX 5090 | RTX 4090 | RTX 3090 |
|---|---|---|---|
| VRAM | 32 GB GDDR7 | 24 GB GDDR6X | 24 GB GDDR6X |
| Bandwidth | 1792 GB/s | 1008 GB/s | 936 GB/s |
| Qwen 3.5 9B Q4 | 115 t/s | 92 t/s | 66 t/s |
| Llama 70B Q4 | 22-28 t/s | 7-10 t/s | 5-8 t/s |
| 2026 price | ≈ 5 700 € (28/09/2026) | ~€1,800 used | ~€750 used |
#2026 used market: smart purchase?
- Buy a used 4090 if
- €1,500–2,000 budget, intensive LLM use, no need for the latest generation. 24 GB of VRAM is a massive argument.
- Prefer a RTX 5090 if
- Budget ≥ €5,700, regular 70B use, need for native FP4, team server. 32 GB changes the equation.
- Prefer a RTX 3090 if
- Budget ≤ €1,000, LLM-only use, speed is secondary. Same VRAM for half the price.
- Prefer a RTX 5070 Ti if
- Budget ~€1,400, usage up to 14B. New with warranty, GDDR7, FP4 support.
#Frequently asked questions
Can the RTX 4090 run Llama 3.3 70B locally?+
How many tokens/sec for Qwen 3.5 9B on RTX 4090?+
Used RTX 4090 or new RTX 5080?+
Does RTX 4090 run hot and use a lot of power with LLMs?+
Can you fine-tune RTX 4090 with LoRA?+
Does the RTX 4090 support FP8?+
What power supply do you need for RTX 4090?+
Prices change quickly: every Monday and Thursday, our tracker records the lowest price for local AI graphics cards, along with the price per GB of VRAM.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.