Home › Catalog › Best LLM on RTX 3090 (24 GB) in 2026 — top used LLM

Best LLM on RTX 3090 (24 GB) in 2026 — top used LLM

◆ Local AI — Your private ChatGPT, free, on your own machine, in an hour · $24 · or all kits $49 →

Ranking updated on 09/10/2026

The RTX 3090 (24 GB GDDR6X, 936 GB/s) is THE best LLM performance-per-dollar option in 2026. ~600-700 € used, 24 GB at full power, Qwen 32B Q5 running smoothly. Stack 2× for 48 GB at ~1300 €.

Offers and alternatives for local AI

RTX 3090 : purchasing alternative available for local AI — GMKtec EVO-X2 64 GB / 1 TB (Ryzen AI Max+ 395) :

A mini PC is a complete machine: check the required memory and software compatibility. It does not replace macOS/MLX or CUDA.

Why this choice? Our complete overview of the GMKtec EVO-X2 64 GB / 1 TB (Ryzen AI Max+ 395) →

Which PC should you choose for your budget? Our picks from €800 to €3,500 →

Affiliate links — BestLLMfor may earn a commission on purchases, at no extra cost to you, which does not influence the ranking (established independently). As an Amazon Associate, BestLLMfor earns from qualifying purchases.

Ranking

1

🇺🇸 Gemma 4 26B-A4B MoE

Google · 26B parameters · Apache 2.0 · 128,000 tokens ctx

MoE variant of Gemma 4. 26B/4B active. Full multimodal (text+image+audio).

Why this ranking MoE variant of Gemma 4. 26B/4B active. Full multimodal (text+image+audio).
ollama run gemma4:26b
On RTX 3090
Q5_K_M
19 GB · 22 tok/s
2

🇨🇳 LLaDA 2.0 Uni 16B

Ant Group / inclusionAI · 16B parameters · Apache 2.0 · 8,192-token context

First open Apache 2.0 dLLM: MoE 16B/1B + 6.2B diffusion decoder. Unified text+vision. Released April 22, 2026.

Why this ranking First open Apache 2.0 dLLM: MoE 16B/1B + 6.2B diffusion decoder. Unified text+vision. Released April 22, 2026.
# HuggingFace : inclusionAI/LLaDA2.0-Uni (Flash Attn 2 + CUDA 12.4 requis)
On RTX 3090
Q5_K_M
22 GB · 60 tok/s
3

🇨🇳 Qwen 3.6 27B

Alibaba · 27B parameters · Apache 2.0 · 262,144-token context

Dense multimodal 27B released April 22, 2026. 262k ctx (1M YaRN). SWE-bench Verified 77.2%.

Why this ranking Dense multimodal 27B released April 22, 2026. 262k ctx (1M YaRN). SWE-bench Verified 77.2%.
ollama run qwen3.6:27b
On RTX 3090
Q5_K_M
19 GB · 13 tok/s
4

🇨🇳 Qwen 3.8 27B

Alibaba · 27B parameters · Apache 2.0 · 262,144-token context

Qwen 3.8 27B: dense multimodal (text + vision), 262k context, ~16 GB Q4 VRAM (18 GB of Ollama weights). Apache 2.0, agentic coding and vision.

Why this ranking Qwen 3.8 27B: dense multimodal (text + vision), 262k context, ~16 GB Q4 VRAM (18 GB of Ollama weights). Apache 2.0, agentic coding and vision.
ollama run qwen3.8:27b
On RTX 3090
Q5_K_M
19 GB · 14 tok/s
5

🇨🇳 GLM 4.7 Flash

Zhipu AI · 31B parameters · MIT · 128,000 tokens ctx

GLM-4.7-Flash (MoE 31B, ~3B active): the best code/VRAM ratio in the 30B class. MIT, 128k ctx, very fast on 3090/4090.

Why this ranking GLM-4.7-Flash (MoE 31B, ~3B active): the best code/VRAM ratio in the 30B class. MIT, 128k ctx, very fast on 3090/4090.
ollama run glm-4.7-flash
On RTX 3090
Q5_K_M
23 GB · 40 tok/s
6

🇺🇸 Gemma 4 31B

Google · 31B parameters · Apache 2.0 · 256,000-token context

Dense 31B multimodal (text+image+audio). 140+ languages, 256k context. #3 open model on Chatbot Arena.

Why this ranking Dense 31B multimodal (text+image+audio). 140+ languages, 256k context. #3 open model on Chatbot Arena.
ollama run gemma4:31b
On RTX 3090
Q5_K_M
22 GB · 12 tok/s
7

🇺🇸 Granite 4.1 30B Instruct

IBM · 30B parameters · Apache 2.0 · 131,072 tokens ctx

Dense 30B Apache 2.0, 12 languages including FR, 131k ctx, GQA 32Q/8KV. OpenAI-compatible tool calling. Released April 29, 2026.

Why this ranking Dense 30B Apache 2.0, 12 languages including FR, 131k ctx, GQA 32Q/8KV. OpenAI-compatible tool calling. Released April 29, 2026.
ollama run granite4.1:30b
On RTX 3090
Q5_K_M
21 GB · 12 tok/s
8

🇺🇸 Nemotron Cascade 2 30B-A3B

NVIDIA · 30B parameters · NVIDIA Open Model License · 128,000 tokens ctx

MoE with 30B/3B active: thinking mode + instruct. Gold medalist at IMO 2025 and IOI 2025. Fast inference thanks to the 3B active parameters, with 30B-level reasoning capabilities. Released April 2026.

Why this ranking MoE with 30B/3B active: thinking mode + instruct. Gold medalist at IMO 2025 and IOI 2025. Fast inference thanks to the 3B active parameters, with 30B-level reasoning capabilities. Released April 2026.
ollama run nemotron-cascade-2
On RTX 3090
Q5_K_M
21 GB · 30 tok/s

Comparison table

Rank Model Params Q4 VRAM Context License On RTX 3090
#1 Gemma 4 26B-A4B MoE 26B 16 GB 128 000 Apache 2.0 22 tok/s · Q5_K_M
#2 LLaDA 2.0 Uni 16B 16B 18 GB 8 192 Apache 2.0 60 tok/s · Q5_K_M
#3 Qwen 3.6 27B 27B 16 GB 262 144 Apache 2.0 13 tok/s · Q5_K_M
#4 Qwen 3.8 27B 27B 16 GB 262 144 Apache 2.0 14 tok/s · Q5_K_M
#5 GLM 4.7 Flash 31B 19 GB 128 000 MIT 40 tok/s · Q5_K_M
#6 Gemma 4 31B 31B 18 GB 256 000 Apache 2.0 12 tok/s · Q5_K_M
#7 Granite 4.1 30B Instruct 30B 17 GB 131 072 Apache 2.0 12 tok/s · Q5_K_M
#8 Nemotron Cascade 2 30B-A3B 30B 17 GB 128 000 NVIDIA Open Model License 30 tok/s · Q5_K_M
The Local AI Kit

Your private, free ChatGPT on your machine in 1 hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

Free memo

Which coding model should you run on YOUR machine?

Get the memo VRAM → best coding model → Ollama command (one screen, copy and paste). Then switch to the Copilote Local kit for a setup that actually works.

The Local Copilot kit — the Ollama + Cline + Aider configs are ready to paste, with tuned Modelfiles, troubleshooting, and lifetime online access →

No spam. Unsubscribe in 1 click. Your data stays with us (never resold).

Ranking methodology

Filter: Q4_K_M ≤ 22 GB. 13–32B bonus (24 GB peak). 936 GB/s GDDR6X = solid Ampere.

Criteria considered:

  • Q4_K_M ≤ 22 GB
  • Best used LLM buy in 2026
  • Qwen 3 32B Q5 smoothly
  • LoRA fine-tuning 7–13B

The scoring is fully transparent: see our methodology for details on VRAM/tokens/sec calculations.

Frequently asked questions

Why is the 3090 the top occasional-use LLM GPU in 2026?

24 GB VRAM (= 4090 and 3090 Ti) at ~€600–700. The new 4070 Ti Super is ~€900 for 16 GB. The 3090 remains unbeatable in €/GB of VRAM for LLMs. See complete guide.

3090 vs 4090: which one should you choose in 2026?

4090 ~2× faster but ~€1,500 new vs. 3090 ~€650 used. If budget is fine, 4090. If you’re being rational, 3090. See RTX 4090.

2× 3090 stack for Llama 70B?

Yes—48 GB total for ~€1,300 used. Llama 70B Q4_K_M (~40 GB) fits with tensor parallelism (vLLM, llama.cpp -tp 2). ~25–35 tok/s. Hard to beat for €/performance with a local 70B.

What is the sweet-spot model on a 3090?

Qwen 3 32B Q5 (~22 GB) at 22-28 tok/s or Mistral Small 24B Q6 (~18 GB) at 30-35 tok/s. For code, Qwen 2.5 Coder 32B Q4 (~17 GB). See code ranking.

Go further

BestLLMfor Kits The reference guide by use case
All kits for life — $49

Prices in euros (€) are French market prices including VAT, as checked by BestLLMfor. US prices differ: the Amazon buttons show the current US price.