Home › Catalog › Best LLM on RTX 3090 Ti (24 GB) in 2026

Best LLM on RTX 3090 Ti (24 GB) in 2026

◆ Local AI — Your private ChatGPT, free, on your own machine, in an hour · $24 · or all kits $49 →

Ranking updated on 09/10/2026

The RTX 3090 Ti (24 GB GDDR6X, 1008 GB/s) is the Ampere flagship. 24 GB + 1 TB/s of bandwidth = the same VRAM capacity as a 4090 at 60% of the new price, ~€700 used.

Offers and alternatives for local AI

RTX 3090 Ti : purchasing alternative available for local AI — GMKtec EVO-X2 64 GB / 1 TB (Ryzen AI Max+ 395) :

A mini PC is a complete machine: check the required memory and software compatibility. It does not replace macOS/MLX or CUDA.

Why this choice? Our complete overview of the GMKtec EVO-X2 64 GB / 1 TB (Ryzen AI Max+ 395) →

Which PC should you choose for your budget? Our picks from €800 to €3,500 →

Affiliate links — BestLLMfor may earn a commission from purchases at no extra cost to you, which does not influence the independently determined ranking. As an Amazon Associate, BestLLMfor earns from qualifying purchases.

Ranking

1

🇺🇸 Gemma 4 26B-A4B MoE

Google · 26B parameters · Apache 2.0 · 128,000 tokens ctx

MoE variant of Gemma 4. 26B/4B active. Full multimodal (text+image+audio).

Why this ranking MoE variant of Gemma 4. 26B/4B active. Full multimodal (text+image+audio).
ollama run gemma4:26b
On RTX 3090 Ti
Q5_K_M
19 GB · 22 tok/s
2

🇨🇳 LLaDA 2.0 Uni 16B

Ant Group / inclusionAI · 16B parameters · Apache 2.0 · 8,192-token context

First open Apache 2.0 dLLM: MoE 16B/1B + 6.2B diffusion decoder. Unified text+vision. Released April 22, 2026.

Why this ranking First open Apache 2.0 dLLM: MoE 16B/1B + 6.2B diffusion decoder. Unified text+vision. Released April 22, 2026.
# HuggingFace : inclusionAI/LLaDA2.0-Uni (Flash Attn 2 + CUDA 12.4 requis)
On RTX 3090 Ti
Q5_K_M
22 GB · 60 tok/s
3

🇨🇳 Qwen 3.6 27B

Alibaba · 27B parameters · Apache 2.0 · 262,144-token context

Dense multimodal 27B released April 22, 2026. 262k ctx (1M YaRN). SWE-bench Verified 77.2%.

Why this ranking Dense multimodal 27B released April 22, 2026. 262k ctx (1M YaRN). SWE-bench Verified 77.2%.
ollama run qwen3.6:27b
On RTX 3090 Ti
Q5_K_M
19 GB · 13 tok/s
4

🇨🇳 Qwen 3.8 27B

Alibaba · 27B parameters · Apache 2.0 · 262,144-token context

Qwen 3.8 27B: dense multimodal (text + vision), 262k context, ~16 GB Q4 VRAM (18 GB of Ollama weights). Apache 2.0, agentic coding and vision.

Why this ranking Qwen 3.8 27B: dense multimodal (text + vision), 262k context, ~16 GB Q4 VRAM (18 GB of Ollama weights). Apache 2.0, agentic coding and vision.
ollama run qwen3.8:27b
On RTX 3090 Ti
Q5_K_M
19 GB · 14 tok/s
5

🇨🇳 GLM 4.7 Flash

Zhipu AI · 31B parameters · MIT · 128,000 tokens ctx

GLM-4.7-Flash (MoE 31B, ~3B active): the best code/VRAM ratio in the 30B class. MIT, 128k ctx, very fast on 3090/4090.

Why this ranking GLM-4.7-Flash (MoE 31B, ~3B active): the best code/VRAM ratio in the 30B class. MIT, 128k ctx, very fast on 3090/4090.
ollama run glm-4.7-flash
On RTX 3090 Ti
Q5_K_M
23 GB · 40 tok/s
6

🇺🇸 Gemma 4 31B

Google · 31B parameters · Apache 2.0 · 256,000-token context

Dense 31B multimodal (text+image+audio). 140+ languages, 256k context. #3 open model on Chatbot Arena.

Why this ranking Dense 31B multimodal (text+image+audio). 140+ languages, 256k context. #3 open model on Chatbot Arena.
ollama run gemma4:31b
On RTX 3090 Ti
Q5_K_M
22 GB · 12 tok/s
7

🇺🇸 Granite 4.1 30B Instruct

IBM · 30B parameters · Apache 2.0 · 131,072 tokens ctx

Dense 30B Apache 2.0, 12 languages including FR, 131k ctx, GQA 32Q/8KV. OpenAI-compatible tool calling. Released April 29, 2026.

Why this ranking Dense 30B Apache 2.0, 12 languages including FR, 131k ctx, GQA 32Q/8KV. OpenAI-compatible tool calling. Released April 29, 2026.
ollama run granite4.1:30b
On RTX 3090 Ti
Q5_K_M
21 GB · 12 tok/s
8

🇺🇸 Nemotron Cascade 2 30B-A3B

NVIDIA · 30B parameters · NVIDIA Open Model License · 128,000 tokens ctx

MoE with 30B/3B active: thinking mode + instruct. Gold medalist at IMO 2025 and IOI 2025. Fast inference thanks to the 3B active parameters, with 30B-level reasoning capabilities. Released April 2026.

Why this ranking MoE with 30B/3B active: thinking mode + instruct. Gold medalist at IMO 2025 and IOI 2025. Fast inference thanks to the 3B active parameters, with 30B-level reasoning capabilities. Released April 2026.
ollama run nemotron-cascade-2
On RTX 3090 Ti
Q5_K_M
21 GB · 30 tok/s

Comparison table

Rank Model Params Q4 VRAM Context License On RTX 3090 Ti
#1 Gemma 4 26B-A4B MoE 26B 16 GB 128 000 Apache 2.0 22 tok/s · Q5_K_M
#2 LLaDA 2.0 Uni 16B 16B 18 GB 8 192 Apache 2.0 60 tok/s · Q5_K_M
#3 Qwen 3.6 27B 27B 16 GB 262 144 Apache 2.0 13 tok/s · Q5_K_M
#4 Qwen 3.8 27B 27B 16 GB 262 144 Apache 2.0 14 tok/s · Q5_K_M
#5 GLM 4.7 Flash 31B 19 GB 128 000 MIT 40 tok/s · Q5_K_M
#6 Gemma 4 31B 31B 18 GB 256 000 Apache 2.0 12 tok/s · Q5_K_M
#7 Granite 4.1 30B Instruct 30B 17 GB 131 072 Apache 2.0 12 tok/s · Q5_K_M
#8 Nemotron Cascade 2 30B-A3B 30B 17 GB 128 000 NVIDIA Open Model License 30 tok/s · Q5_K_M
The Local AI Kit

Your private, free ChatGPT on your machine in 1 hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

Free memo

Which coding model should you run on YOUR machine?

Get the memo VRAM → best coding model → Ollama command (one screen, copy and paste). Then switch to the Copilote Local kit for a setup that actually works.

The Local Copilot kit — the Ollama + Cline + Aider configs are ready to paste, with tuned Modelfiles, troubleshooting, and lifetime online access →

No spam. Unsubscribe in 1 click. Your data stays with us (never resold).

Ranking methodology

Filter: Q4_K_M ≤ 22 GB. Bonus 13-32B (peak 24 GB) and 7-32B. Record Ampere bandwidth of 1008 GB/s.

Criteria considered:

  • Q4_K_M ≤ 22 GB
  • Qwen 3 32B Q5 with comfortable headroom
  • LoRA fine-tuning 7–13B
  • 24 GB + 1 TB/s

The scoring is fully transparent: see our methodology for details on VRAM/tokens/sec calculations.

Frequently asked questions

RTX 3090 Ti vs 3090?

Even at 24 GB. 3090 Ti = +7% CUDA cores + GDDR6X 1008 GB/s vs. 3090 GDDR6X 936 GB/s. Difference ~5–8% for LLMs. 3090 is often the better used-market deal. See RTX 3090.

3090 Ti vs. 4090?

Same 24 GB. 4090 = 1008 GB/s too, plus 16384 CUDA cores vs. 10752 on the 3090 Ti. ~40-50% faster for LLMs. If buying new, get the 4090. Used, ~€700 vs. ~€1100, the 3090 Ti is excellent. See RTX 4090.

Llama 70B on a 3090 Ti?

Q3_K_M (~32 GB) doesn't fit by itself. Q2_K (~24 GB) just fits, but with degraded quality. For comfortable 70B use, 2× 3090 Ti or RTX 5090 32 GB. See RTX 5090.

Used 2× 3090 Ti setup?

Excellent: 48 GB of split VRAM for ~€1,400 total. Llama 70B Q4 (~40 GB) fits and reaches ~30 tok/s via tensor parallelism. Hard to beat for LLM price/performance in 2026.

Go further

QuelLLM Kits The reference guide by use case
All kits for life — $49

Prices in euros (€) are French market prices including VAT, checked by QuelLLM. US prices differ: the Amazon buttons show the current US price.