Home › Catalog › Best LLM on Radeon RX 7900 XT (20 GB) in 2026

Best LLM on Radeon RX 7900 XT (20 GB) in 2026

◆ Local AI — Your private ChatGPT, free, on your own machine, in an hour · $24 · or all kits $49 →

Ranking updated on 09/10/2026

The Radeon RX 7900 XT (20 GB GDDR6, 800 GB/s) is the younger sibling of the 7900 XTX. Its unusual 20 GB makes comfortable 24–32B models in Q4 possible. ~€650 new, excellent value.

Offers and alternatives for local AI

Radeon RX 7900 XT : purchasing alternative available for local AI — Radeon RX 9070 XT 16 GB :

Why this choice? Our complete guide to Radeon RX 9070 XT 16 GB →

Which PC should you choose for your budget? Our picks from €800 to €3,500 →

Affiliate links — BestLLMfor may earn a commission from purchases at no extra cost to you, which does not influence the independently determined ranking. As an Amazon Associate, BestLLMfor earns from qualifying purchases.

Ranking

1

🇨🇳 LLaDA 2.0 Uni 16B

Ant Group / inclusionAI · 16B parameters · Apache 2.0 · 8,192-token context

First open Apache 2.0 dLLM: MoE 16B/1B + 6.2B diffusion decoder. Unified text+vision. Released April 22, 2026.

Why this ranking First open Apache 2.0 dLLM: MoE 16B/1B + 6.2B diffusion decoder. Unified text+vision. Released April 22, 2026.
# HuggingFace : inclusionAI/LLaDA2.0-Uni (Flash Attn 2 + CUDA 12.4 requis)
On Radeon RX 7900 XT
Q4_K_M
18 GB · 60 tok/s
2

🇫🇷 Devstral Small 2 24B

Mistral AI · 24B parameters · Apache 2.0 · 256,000-token context

24B coding specialist, Apache 2.0. 72.2% SWE-Bench. 256k ctx, FR lab.

Why this ranking 24B coding specialist, Apache 2.0. 72.2% SWE-Bench. 256k ctx, FR lab.
ollama run devstral-small2:24b
On Radeon RX 7900 XT
Q5_K_M
17 GB · 15 tok/s
3

🇺🇸 Gemma 4 26B-A4B MoE

Google · 26B parameters · Apache 2.0 · 128,000 tokens ctx

MoE variant of Gemma 4. 26B/4B active. Full multimodal (text+image+audio).

Why this ranking MoE variant of Gemma 4. 26B/4B active. Full multimodal (text+image+audio).
ollama run gemma4:26b
On Radeon RX 7900 XT
Q5_K_M
19 GB · 22 tok/s
4

🇨🇳 Qwen 3.6 27B

Alibaba · 27B parameters · Apache 2.0 · 262,144-token context

Dense multimodal 27B released April 22, 2026. 262k ctx (1M YaRN). SWE-bench Verified 77.2%.

Why this ranking Dense multimodal 27B released April 22, 2026. 262k ctx (1M YaRN). SWE-bench Verified 77.2%.
ollama run qwen3.6:27b
On Radeon RX 7900 XT
Q5_K_M
19 GB · 13 tok/s
5

🇨🇳 Qwen 3.8 27B

Alibaba · 27B parameters · Apache 2.0 · 262,144-token context

Qwen 3.8 27B: dense multimodal (text + vision), 262k context, ~16 GB Q4 VRAM (18 GB of Ollama weights). Apache 2.0, agentic coding and vision.

Why this ranking Qwen 3.8 27B: dense multimodal (text + vision), 262k context, ~16 GB Q4 VRAM (18 GB of Ollama weights). Apache 2.0, agentic coding and vision.
ollama run qwen3.8:27b
On Radeon RX 7900 XT
Q5_K_M
19 GB · 14 tok/s
6

🇫🇷 Magistral Small 24B

Mistral AI · 24B parameters · Apache 2.0 · 128,000 tokens ctx

First open Mistral reasoner. AIME24 70.7%. Based on Small 3.1 + CoT training.

Why this ranking First open Mistral reasoner. AIME24 70.7%. Based on Small 3.1 + CoT training.
ollama run magistral:24b
On Radeon RX 7900 XT
Q5_K_M
17 GB · 15 tok/s
7

🇺🇸 Gemma 4 31B

Google · 31B parameters · Apache 2.0 · 256,000-token context

Dense 31B multimodal (text+image+audio). 140+ languages, 256k context. #3 open model on Chatbot Arena.

Why this ranking Dense 31B multimodal (text+image+audio). 140+ languages, 256k context. #3 open model on Chatbot Arena.
ollama run gemma4:31b
On Radeon RX 7900 XT
Q4_K_M
18 GB · 12 tok/s

Comparison table

Rank Model Params Q4 VRAM Context License On Radeon RX 7900 XT
#1 LLaDA 2.0 Uni 16B 16B 18 GB 8 192 Apache 2.0 60 tok/s · Q4_K_M
#2 Devstral Small 2 24B 24B 14 GB 256 000 Apache 2.0 15 tok/s · Q5_K_M
#3 Gemma 4 26B-A4B MoE 26B 16 GB 128 000 Apache 2.0 22 tok/s · Q5_K_M
#4 Qwen 3.6 27B 27B 16 GB 262 144 Apache 2.0 13 tok/s · Q5_K_M
#5 Qwen 3.8 27B 27B 16 GB 262 144 Apache 2.0 14 tok/s · Q5_K_M
#6 Magistral Small 24B 24B 14 GB 128 000 Apache 2.0 15 tok/s · Q5_K_M
#7 Gemma 4 31B 31B 18 GB 256 000 Apache 2.0 12 tok/s · Q4_K_M
The Local AI Kit

Your private, free ChatGPT on your machine in 1 hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

Free memo

Which coding model should you run on YOUR machine?

Get the memo VRAM → best coding model → Ollama command (one screen, copy and paste). Then switch to the Copilote Local kit for a setup that actually works.

The Local Copilot kit — the Ollama + Cline + Aider configs are ready to paste, with tuned Modelfiles, troubleshooting, and lifetime online access →

No spam. Unsubscribe in 1 click. Your data stays with us (never resold).

Ranking methodology

Filter: Q4_K_M ≤ 18 GB. 13-32B bonus (20 GB peak). 800 GB/s + ROCm 6.

Criteria considered:

  • Q4_K_M ≤ 18 GB
  • ROCm 6 compatible
  • Mistral Small 24B Q5
  • Unusual 20 GB

The scoring is fully transparent: see our methodology for details on VRAM/tokens/sec calculations.

Frequently asked questions

RX 7900 XT versus 7900 XTX?

XT = 20 GB + 800 GB/s. XTX = 24 GB + 960 GB/s. For 24–32B Q4, both work. XTX is preferable for fine-tuning or 32B Q5. See RX 7900 XTX.

Why 20 GB instead of 16 or 24?

AMD’s marketing choice to position it between the 7900 XTX and the previous tier. For LLMs, it is the sweet spot: Mistral Small 24B Q4 (~13 GB) + 32k context + cache = ~18 GB.

ROCm on a 7900 XT: easy?

Yes, since ROCm 6 (2024). Ollama has native support via -gpu rocm. llama.cpp does too. Simpler than it was 2 years ago. See guide.

Sweet spot LLM for the 7900 XT?

Mistral Small 24B Q5 (~17 GB) at 25 tok/s, Qwen 3 32B Q4 (~17 GB) at 22 tok/s. Excellent for code + chat.

Go further

QuelLLM Kits The reference guide by use case
All kits for life — $49

Prices in euros (€) are French market prices including VAT, checked by QuelLLM. US prices differ: the Amazon buttons show the current US price.