Home › Catalog › Best LLM on RTX 4090 (24 GB) in 2026

Best LLM on RTX 4090 (24 GB) in 2026

◆ Local AI — Your private ChatGPT, free, on your own machine, in an hour · $24 · or all kits $49 →

Ranking updated on 09/10/2026

The RTX 4090 (24 GB VRAM, Ada Lovelace architecture) is the mainstream reference for LLM inference in 2026. Here are the models that make the most of it: top-quality Q4/Q5 models that fit in 24 GB, with comfortable throughput (30+ tokens/sec).

Offers and alternatives for local AI

RTX 4090: purchasing alternative available for local AI — GMKtec EVO-X2 64 GB / 1 TB (Ryzen AI Max+ 395):

A mini PC is a complete machine: check the required memory and software compatibility. It does not replace macOS/MLX or CUDA.

Why this choice? Our complete overview of the GMKtec EVO-X2 64 GB / 1 TB (Ryzen AI Max+ 395) →

Which PC should you choose for your budget? Our picks from €800 to €3,500 →

Affiliate links — BestLLMfor may earn a commission on purchases, at no extra cost to you, which does not influence the ranking (established independently). As an Amazon Associate, BestLLMfor earns from qualifying purchases.

Ranking

1

🇺🇸 Laguna XS.2

Poolside · 33B parameters · Apache 2.0 · 131,072 tokens ctx

MoE 33B/3B active parameters, Apache 2.0, specializing in agentic coding. 68.2% SWE-Bench Verified, 128k ctx. Runs on a 36 GB Mac. Released April 28, 2026.

Why this ranking Fits in Q5_K_M (~23 GB out of 24 GB available). 33B parameters, 131,072-token context.
ollama run laguna-xs.2
On RTX 4090
Q5_K_M
23 GB · 100 tok/s
2

🇺🇸 Gemma 4 26B-A4B MoE

Google · 26B parameters · Apache 2.0 · 128,000 tokens ctx

MoE variant of Gemma 4. 26B/4B active. Full multimodal (text+image+audio).

Why this ranking Fits in Q5_K_M (~19 GB out of 24 GB available). 26B parameters, 128,000-token context.
ollama run gemma4:26b
On RTX 4090
Q5_K_M
19 GB · 60 tok/s
3

🇨🇳 LLaDA 2.0 Uni 16B

Ant Group / inclusionAI · 16B parameters · Apache 2.0 · 8,192-token context

First open Apache 2.0 dLLM: MoE 16B/1B + 6.2B diffusion decoder. Unified text+vision. Released April 22, 2026.

Why this ranking Fits in Q5_K_M (~22 GB out of 24 GB available). 16B parameters, 8,192-token context.
# HuggingFace : inclusionAI/LLaDA2.0-Uni (Flash Attn 2 + CUDA 12.4 requis)
On RTX 4090
Q5_K_M
22 GB · 130 tok/s
4

🇨🇳 Qwen 3.6 27B

Alibaba · 27B parameters · Apache 2.0 · 262,144-token context

Dense multimodal 27B released April 22, 2026. 262k ctx (1M YaRN). SWE-bench Verified 77.2%.

Why this ranking Fits in Q5_K_M (~19 GB out of 24 GB available). 27B parameters, 262,144-token context.
ollama run qwen3.6:27b
On RTX 4090
Q5_K_M
19 GB · 32 tok/s
5

🇨🇳 Qwen 3.8 27B

Alibaba · 27B parameters · Apache 2.0 · 262,144-token context

Qwen 3.8 27B: dense multimodal (text + vision), 262k context, ~16 GB Q4 VRAM (18 GB of Ollama weights). Apache 2.0, agentic coding and vision.

Why this ranking Fits in Q5_K_M (~19 GB out of 24 GB available). 27B parameters, 262,144-token context.
ollama run qwen3.8:27b
On RTX 4090
Q5_K_M
19 GB · 22 tok/s
6

🇨🇳 GLM 4.7 Flash

Zhipu AI · 31B parameters · MIT · 128,000 tokens ctx

GLM-4.7-Flash (MoE 31B, ~3B active): the best code/VRAM ratio in the 30B class. MIT, 128k ctx, very fast on 3090/4090.

Why this ranking Fits in Q5_K_M (~23 GB out of 24 GB available). 31B parameters, 128,000-token context.
ollama run glm-4.7-flash
On RTX 4090
Q5_K_M
23 GB · 100 tok/s
7

🇺🇸 Gemma 4 31B

Google · 31B parameters · Apache 2.0 · 256,000-token context

Dense 31B multimodal (text+image+audio). 140+ languages, 256k context. #3 open model on Chatbot Arena.

Why this ranking Fits in Q5_K_M (~22 GB out of 24 GB available). 31B parameters, 256,000-token context.
ollama run gemma4:31b
On RTX 4090
Q5_K_M
22 GB · 30 tok/s

Comparison table

Rank Model Params Q4 VRAM Context License On RTX 4090
#1 Laguna XS.2 33B 19 GB 131 072 Apache 2.0 100 tok/s · Q5_K_M
#2 Gemma 4 26B-A4B MoE 26B 16 GB 128 000 Apache 2.0 60 tok/s · Q5_K_M
#3 LLaDA 2.0 Uni 16B 16B 18 GB 8 192 Apache 2.0 130 tok/s · Q5_K_M
#4 Qwen 3.6 27B 27B 16 GB 262 144 Apache 2.0 32 tok/s · Q5_K_M
#5 Qwen 3.8 27B 27B 16 GB 262 144 Apache 2.0 22 tok/s · Q5_K_M
#6 GLM 4.7 Flash 31B 19 GB 128 000 MIT 100 tok/s · Q5_K_M
#7 Gemma 4 31B 31B 18 GB 256 000 Apache 2.0 30 tok/s · Q5_K_M
The Local AI Kit

Your private, free ChatGPT on your machine in 1 hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

Free memo

Which coding model should you run on YOUR machine?

Get the memo VRAM → best coding model → Ollama command (one screen, copy and paste). Then switch to the Copilote Local kit for a setup that actually works.

The Local Copilot kit — the Ollama + Cline + Aider configs are ready to paste, with tuned Modelfiles, troubleshooting, and lifetime online access →

No spam. Unsubscribe in 1 click. Your data stays with us (never resold).

Ranking methodology

We keep models that fit within 24 GB in Q4_K_M and use at least 40% of the VRAM (otherwise a 7B is sufficient). Bonus points go to models whose VRAM fit is between 60% and 95%—the quality/throughput sweet spot.

Criteria considered:

  • Fits in 24 GB in Q4_K_M
  • Takes advantage of VRAM (> 60%)
  • Throughput ≥ 30 tokens/sec
  • Quality > 7B

The scoring is fully transparent: see our methodology for details on VRAM/tokens/sec calculations.

Frequently asked questions

Can you run a 70B on RTX 4090?

In Q4_K_M, a 70B model requires ~40 GB of VRAM—too much for a single 4090. You must either drop to Q2/Q3 (quality loss), offload to CPU RAM (very slow), or add a 2nd card. For a true 70B model, aim for 2× RTX 4090 or a 5090 + DDR5.

Which quantization should you choose on RTX 4090?

Q5_K_M is the sweet spot (less than 1% loss versus FP16 according to benchmarks). Q8 is clearly better than Q5 but uses 50% more VRAM. Use Q4 only if you want a large model that won't fit in Q5.

Mistral Small 3.1 24B or Qwen 2.5 32B on a 4090?

View the comparison. Mistral Small 3.1 is faster (24B < 32B) and better in French. Qwen 2.5 32B is more capable for general tasks and code.

Which inference engine for RTX 4090?

For interactive chat: Ollama (simple) or llama.cpp (maximum control). For server throughput: vLLM or ExLlamaV2. The gain can reach 2–3× on vLLM in batch.

Go further

BestLLMfor Kits The reference guide by use case
All kits for life — $49

Prices in euros (€) are French market prices including VAT, as checked by BestLLMfor. US prices differ: the Amazon buttons show the current US price.