Intermediate 12 minRTX 40

Which LLM on RTX 4090 (24 GB) ?

The RTX 4090 (October 2022) remains the benchmark for mainstream local AI in 2026: 24 GB of GDDR6X VRAM, 1008 GB/s of bandwidth, and 16,384 CUDA Ada Lovelace cores. It runs Qwen 3.5 9B at over 90 tok/s, a Qwen 3.8 27B Q4 comfortably, and even a Llama 70B Q4 with light offload. Now selling for €1700–2000 used (ex-LHR stock), it offers a better performance-per-euro ratio than a new 5080 for pure LLM use. This guide details all the models it can run and approximate throughput figures.

Choosing a machine? Our picks by budget →

By Mohamed Meguedmi·Update 2026-08-27·Tested on Windows, macOS, and Linux
Recommended hardware

Buying alternative for this guide: GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395).

A mini PC is a complete machine: check the available memory and engine compatibility. It does not replace macOS/MLX or CUDA.

Why this choice? Our complete guide on GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395) →

Compare all options by budget, from €800 to €3,500 →

On the go: which laptop for local AI →

Affiliate links — commission possible at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.

#The RTX 4090 in 2026, still the queen

Architecture
Ada Lovelace (AD102), manufactured on TSMC 4N. 76 billion transistors.
VRAM
24 GB GDDR6X at 21 Gbps, 384-bit bus. 1008 GB/s bandwidth.
CUDA cores
16,384. 4th-gen Tensor Cores with native FP8. 3rd-gen RT Cores.
TDP
450 W. An 850 W PSU is recommended.
2026 price
1700–2000 € in good used condition.
→
Why it still matters in 2026
24 GB = the magic threshold that makes Llama 3.3 70B Q2_K possible (26 GB with 2 GB offload), Qwen 3.8 27B in Q8, Qwen3-Coder 30B-A3B. The 5080 (16 GB new) cannot do it; the 5090 (32 GB new, ≈ €5,700) costs about 3× as much. In pure VRAM/€ value, the used 4090 dominates.

#1. Compatible models 24 GB

LLM models — RTX 4090 24 GB
ModelQuantVRAMTokens/secUsage
gpt-oss 20BQ4_K_M13 GB117-143Instant chat
Devstral Small 2 24BQ4_K_M14 GB36-44Chat + RAG FR
Qwen 3.6 27BQ4_K_M16 GB29-35Pro code assistant
Qwen 3.8 27BQ4_K_M16 GB20-24Quality reasoning
Mistral Small 24BQ6_K20 GB32-38Versatile sweet spot
GLM 4.7 FlashQ4_K_M19 GB90-110Advanced reasoning
Granite 4.1 30B InstructQ4_K_M17 GB27-33Complex multi-file code
Gemma 4 26B-A4B MoEQ5_K_M19 GB54-662 GB RAM offload, usable
Gemma 4 31BQ5_K_M22 GB27-33All in VRAM, degraded quality
i
Llama 70B: which quantization should you choose?
Q4_K_M (42 GB): heavy offload, ~6 tok/s, not very comfortable. Q2_K (26 GB): 2 GB to offload to RAM, ~10 tok/s, decent quality. IQ2_XS (21 GB): everything in VRAM, ~20 tok/s but noticeably degraded quality. For a comfortable 70B model locally, target a 5090 (32 GB) or a Mac Studio.

#2. Installation & CUDA

  1. 01
    NVIDIA 550+ drivers
    CUDA 12.4+. For a new 4090, GeForce Experience handles it. Linux: nvidia-driver-565 or later.
  2. 02
    Ollama
    Standard installer from ollama.com. CUDA GPU auto-detection.
  3. 03
    First 24 GB model test
    ollama run qwen3.8:27b — profite de la VRAM dispo et met en valeur la carte.
Complete Linux setup
# Ubuntu 24.04
sudo apt install nvidia-driver-565
sudo reboot

curl -fsSL https://ollama.com/install.sh | sh
ollama run qwen3.8:27b

#3. Tokens/sec benchmarks

RTX 4090 24 GB — tokens/sec orders of magnitude
ModelQ4_K_MQ5_K_MQ6_KQ8_0
Qwen 3.5 4B135 t/s122 t/s110 t/s90 t/s
Granite 4.2 8B100 t/s92 t/s82 t/s66 t/s
Qwen 3.5 9B92 t/s85 t/s76 t/s62 t/s
Gemma 4 12B70 t/s64 t/s57 t/s48 t/s
Mistral Small 24B42 t/s38 t/s34 t/s28 t/s
Qwen 3.8 27B32 t/s28 t/s25 t/s—

#4. Ada optimizations

Flash Attention 2
Enabled by default. FA3 is coming in future llama.cpp versions.
KV cache Q8
OLLAMA_KV_CACHE_TYPE=q8_0. Enables a 32k context on 14B in ~15 GB, leaving 9 GB for other processes.
Power limit
nvidia-smi -pl 350 (starting at 350 W). LLM: -2% speed, -25% heat and noise. Sustainable 24/7 configuration.
FP8 via TensorRT-LLM
Ada supports native FP8. Via TensorRT-LLM, +40% batch throughput on Qwen 3.5 9B. Especially useful for multi-user serving.

#5. 4090 vs 5090 vs 3090

The three mainstream NVIDIA flagships
CriterionRTX 5090RTX 4090RTX 3090
VRAM32 GB GDDR724 GB GDDR6X24 GB GDDR6X
Bandwidth1792 GB/s1008 GB/s936 GB/s
Qwen 3.5 9B Q4115 t/s92 t/s66 t/s
Llama 70B Q422-28 t/s7-10 t/s5-8 t/s
2026 price≈ 5 700 € (28/09/2026)~€1,800 used~€750 used
→
Best 24 GB/€ value: used 3090
At €750 vs. €1,800, the 3090 offers the same 24 GB of VRAM. About 30% slower in raw speed. If your priority is VRAM (24–32B models, 14B+ QLoRA) rather than maximum speed, a used 3090 remains the best LLM value in all of NVIDIA.

#2026 used market: smart purchase?

Buy a used 4090 if
€1,500–2,000 budget, intensive LLM use, no need for the latest generation. 24 GB of VRAM is a massive argument.
Prefer a RTX 5090 if
Budget ≥ €5,700, regular 70B use, need for native FP4, team server. 32 GB changes the equation.
Prefer a RTX 3090 if
Budget ≤ €1,000, LLM-only use, speed is secondary. Same VRAM for half the price.
Prefer a RTX 5070 Ti if
Budget ~€1,400, usage up to 14B. New with warranty, GDDR7, FP4 support.

#Frequently asked questions

Can the RTX 4090 run Llama 3.3 70B locally?+
Yes, but with trade-offs. Q4_K_M (42 GB): heavy offloading → 6-8 tok/s (unusable). Q2_K (26 GB): 2 GB in RAM → 10-12 tok/s (acceptable). IQ2_XS (21 GB): all in VRAM → 20 tok/s but degraded quality. For a genuinely comfortable large model, prefer a 5090 (32 GB)—or stick with a Qwen 3.8 27B that fits entirely in VRAM.
How many tokens/sec for Qwen 3.5 9B on RTX 4090?+
About 90 tokens/second in Q4_K_M, 62 tok/s in Q8_0. More than sufficient for any chat use case, plus the model’s 256k context and vision.
Used RTX 4090 or new RTX 5080?+
Used 4090 (€1,800): 24 GB VRAM, 10% slower than a 5080 on 7B but enables 70B with offloading and 32B in Q6. New 5080 (≈ €1,500 to €1,850): 16 GB, faster on small models, 3-year warranty, FP4 support. For pure LLM use, the 4090 wins because of its VRAM. For performance plus warranty, the 5080.
Does RTX 4090 run hot and use a lot of power with LLMs?+
Less than when gaming. Sustained Qwen 3.5 9B inference: 250–320 W (vs. 450 W TDP), GPU at 65–72 °C. A properly ventilated case is enough. For 24/7 use, undervolting to 350 W reduces heat and noise without significant LLM performance loss.
Can you fine-tune RTX 4090 with LoRA?+
Yes, excellent. LoRA on Qwen 3.5 9B: 12 GB, comfortable with batch 4 and 4096 context. QLoRA on Mistral Small 24B: ~18 GB, possible. QLoRA on a 70B: tight, but feasible with unsloth + batch 1. For standard 70B LoRA, you need 2× 4090s or a 5090.
Does the RTX 4090 support FP8?+
Yes, Ada Lovelace 4th-generation Tensor Cores support FP8 natively. It is used by TensorRT-LLM (a +30-40% gain vs FP16), vLLM (0.5+), and partially by llama.cpp. For mainstream Ollama, stick with Q4-Q8. For batch serving, FP8 is worth it.
What power supply do you need for RTX 4090?+
850 W minimum, quality unit (ATX 3.0 recommended but not required). In LLM use, average consumption is 250–320 W. In gaming or benchmarks, 450 W. High-end CPU (9950X, 14900K): 1000 W provides more headroom.

Prices change quickly: every Monday and Thursday, our tracker records the lowest price for local AI graphics cards, along with the price per GB of VRAM.

Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.

Prices in euros (€) are French market prices including VAT, as checked by BestLLMfor. US prices differ: the Amazon buttons show the current US price.