Home › Catalog › Best LLM on RTX 5090 (32 GB) in 2026

Best LLM on RTX 5090 (32 GB) in 2026

◆ Local AI — Your private ChatGPT, free, on your own machine, in an hour · $24 · or all kits $49 →

Ranking updated on 09/10/2026

The RTX 5090 (Blackwell, 32 GB GDDR7, 1792 GB/s) is the first consumer GPU to exceed 24 GB. Llama 70B Q4_K_M fits with 12 GB of headroom for context. The definitive local reference for 2026.

Offers and alternatives for local AI

RTX 5090 : purchasing alternative available for local AI — GMKtec EVO-X2 64 GB / 1 TB (Ryzen AI Max+ 395) :

A mini PC is a complete machine: check the required memory and software compatibility. It does not replace macOS/MLX or CUDA.

Why this choice? Our complete overview of the GMKtec EVO-X2 64 GB / 1 TB (Ryzen AI Max+ 395) →

Which PC should you choose for your budget? Our picks from €800 to €3,500 →

Affiliate links — BestLLMfor may earn a commission on purchases, at no extra cost to you, which does not influence the ranking (established independently). As an Amazon Associate, BestLLMfor earns from qualifying purchases.

Ranking

1

🇨🇳 GLM 4.7 Flash

Zhipu AI · 31B parameters · MIT · 128,000 tokens ctx

GLM-4.7-Flash (MoE 31B, ~3B active): the best code/VRAM ratio in the 30B class. MIT, 128k ctx, very fast on 3090/4090.

Why this ranking GLM-4.7-Flash (MoE 31B, ~3B active): the best code/VRAM ratio in the 30B class. MIT, 128k ctx, very fast on 3090/4090.
ollama run glm-4.7-flash
On RTX 5090
Q5_K_M
23 GB · 100 tok/s
2

🇺🇸 Laguna XS.2

Poolside · 33B parameters · Apache 2.0 · 131,072 tokens ctx

MoE 33B/3B active parameters, Apache 2.0, specializing in agentic coding. 68.2% SWE-Bench Verified, 128k ctx. Runs on a 36 GB Mac. Released April 28, 2026.

Why this ranking MoE 33B/3B active parameters, Apache 2.0, specializing in agentic coding. 68.2% SWE-Bench Verified, 128k ctx. Runs on a 36 GB Mac. Released April 28, 2026.
ollama run laguna-xs.2
On RTX 5090
Q5_K_M
23 GB · 100 tok/s
3

🇺🇸 Nemotron Cascade 2 30B-A3B

NVIDIA · 30B parameters · NVIDIA Open Model License · 128,000 tokens ctx

MoE with 30B/3B active: thinking mode + instruct. Gold medalist at IMO 2025 and IOI 2025. Fast inference thanks to the 3B active parameters, with 30B-level reasoning capabilities. Released April 2026.

Why this ranking MoE with 30B/3B active: thinking mode + instruct. Gold medalist at IMO 2025 and IOI 2025. Fast inference thanks to the 3B active parameters, with 30B-level reasoning capabilities. Released April 2026.
ollama run nemotron-cascade-2
On RTX 5090
Q8
32 GB · 80 tok/s
4

🇺🇸 Granite 4.0 H-Small 32B-A9B

IBM · 32B parameters · Apache 2.0 · 128,000 tokens ctx

Mamba-2 + MoE 32B/9B hybrid. ~70% less RAM in long contexts. Apache 2.0.

Why this ranking Mamba-2 + MoE 32B/9B hybrid. ~70% less RAM in long contexts. Apache 2.0.
ollama run granite4:small-h
On RTX 5090
Q5_K_M
23 GB · 75 tok/s
5

🇨🇳 Qwen 3 30B-A3B

Alibaba · 30B parameters · Apache 2.0 · 131,072 tokens ctx

MoE 30B/3B active hybrid thinking. MMLU 81.4, AIME24 80.4. 100+ languages.

Why this ranking MoE 30B/3B active hybrid thinking. MMLU 81.4, AIME24 80.4. 100+ languages.
ollama run qwen3:30b-a3b
On RTX 5090
Q5_K_M
23 GB · 100 tok/s
6

🇺🇸 Nemotron 3 Nano 30B-A3B

NVIDIA · 30B parameters · NVIDIA Open Model License · 128,000 tokens ctx

MoE 30B/3.5B active NVIDIA: chat, code, reasoning. 128k context, 3.5B speed with 30B capabilities. April 2026 release.

Why this ranking MoE 30B/3.5B active NVIDIA: chat, code, reasoning. 128k context, 3.5B speed with 30B capabilities. April 2026 release.
ollama run nemotron-3-nano
On RTX 5090
Q8
32 GB · 80 tok/s
7

🇨🇳 Qwen3-Coder 30B-A3B

Alibaba · 30B parameters · Apache 2.0 · 262,144-token context

MoE 30B (3.3B active parameters) specialized in agentic coding. Very fast locally, native 256k ctx, the benchmark for 16–24 GB via Ollama.

Why this ranking MoE 30B (3.3B active parameters) specialized in agentic coding. Very fast locally, native 256k ctx, the benchmark for 16–24 GB via Ollama.
ollama run qwen3-coder:30b
On RTX 5090
Q5_K_M
23 GB · 100 tok/s
8

🇨🇳 Qwen 3 VL 30B-A3B

Alibaba · 30B parameters · Apache 2.0 · 262,144-token context

Vision MoE with 30B/3B active. Vision sweet spot Qwen 3. 256k ctx.

Why this ranking Vision MoE with 30B/3B active. Vision sweet spot Qwen 3. 256k ctx.
ollama run qwen3-vl:30b
On RTX 5090
Q5_K_M
23 GB · 100 tok/s

Comparison table

Rank Model Params Q4 VRAM Context License On RTX 5090
#1 GLM 4.7 Flash 31B 19 GB 128 000 MIT 100 tok/s · Q5_K_M
#2 Laguna XS.2 33B 19 GB 131 072 Apache 2.0 100 tok/s · Q5_K_M
#3 Nemotron Cascade 2 30B-A3B 30B 17 GB 128 000 NVIDIA Open Model License 80 tok/s · Q8
#4 Granite 4.0 H-Small 32B-A9B 32B 19 GB 128 000 Apache 2.0 75 tok/s · Q5_K_M
#5 Qwen 3 30B-A3B 30B 19 GB 131 072 Apache 2.0 100 tok/s · Q5_K_M
#6 Nemotron 3 Nano 30B-A3B 30B 17 GB 128 000 NVIDIA Open Model License 80 tok/s · Q8
#7 Qwen3-Coder 30B-A3B 30B 19 GB 262 144 Apache 2.0 100 tok/s · Q5_K_M
#8 Qwen 3 VL 30B-A3B 30B 19 GB 262 144 Apache 2.0 100 tok/s · Q5_K_M
The Local AI Kit

Your private, free ChatGPT on your machine in 1 hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

Free memo

Which coding model should you run on YOUR machine?

Get the memo VRAM → best coding model → Ollama command (one screen, copy and paste). Then switch to the Copilote Local kit for a setup that actually works.

The Local Copilot kit — the Ollama + Cline + Aider configs are ready to paste, with tuned Modelfiles, troubleshooting, and lifetime online access →

No spam. Unsubscribe in 1 click. Your data stays with us (never resold).

Ranking methodology

Filter: models whose Q4_K_M fits under 30 GB (leaves 2 GB for context). Bonus: 30–70B (5090 peak) and 100B MoE (32 GB unlocks DBRX, Mixtral 8x22B). 1792 GB/s = record consumer throughput.

Criteria considered:

  • Q4_K_M ≤ 30 GB
  • Leverages 1792 GB/s of bandwidth
  • Smooth 70B Q4/Q5
  • Accessible 100B MoE

The scoring is fully transparent: see our methodology for details on VRAM/tokens/sec calculations.

Frequently asked questions

RTX 5090 32 GB: Llama 70B smoothly?

Yes — Llama 3.3 70B Q4_K_M (~40 GB) does NOT fit by itself; Q3_K_M (~30 GB) runs at 35-45 tok/s. Q5_K_M (~48 GB) requires partial CPU offload. For 70B Q4 without compromises, target 2× RTX 4090/5090 or a Mac Studio with 96+ GB.

RTX 5090 vs. 2× RTX 4090?

5090 = 32 GB monolithic + 1792 GB/s. 2× 4090 = 48 GB (split) + 1008 GB/s per card. For 70B Q4 (~40 GB), 2× 4090 wins. For 30–32B Q5 + long context, 5090 is simpler (no split overhead). See RTX 4090.

Which quantization is optimal on a 5090?

Q5_K_M for 30B (~22 GB) or Q4 for 70B (~40 GB partial offload). Q8 for 13-14B (Qwen 3 14B ~15 GB) for maximum quality. Q6_K is an excellent compromise for 32B (~25 GB).

MoE on RTX 5090?

Excellent: Mixtral 8x22B Q4 (~80 GB) won't fit, but 8x7B Q4 (~28 GB) runs at 80+ tok/s. Qwen 3 30B-A3B Q8 (~32 GB) is smooth too. See agent/MoE ranking.

Go further

BestLLMfor Kits The reference guide by use case
All kits for life — $49

Prices in euros (€) are French market prices including VAT, as checked by BestLLMfor. US prices differ: the Amazon buttons show the current US price.