🇨🇳 LLaDA 2.0 Uni 16B
First open Apache 2.0 dLLM: MoE 16B/1B + 6.2B diffusion decoder. Unified text+vision. Released April 22, 2026.
# HuggingFace : inclusionAI/LLaDA2.0-Uni (Flash Attn 2 + CUDA 12.4 requis)
Ranking updated on 09/10/2026
The Radeon RX 7900 XT (20 GB GDDR6, 800 GB/s) is the younger sibling of the 7900 XTX. Its unusual 20 GB makes comfortable 24–32B models in Q4 possible. ~€650 new, excellent value.
Radeon RX 7900 XT : purchasing alternative available for local AI — Radeon RX 9070 XT 16 GB :
Why this choice? Our complete guide to Radeon RX 9070 XT 16 GB →
Which PC should you choose for your budget? Our picks from €800 to €3,500 →
Affiliate links — BestLLMfor may earn a commission from purchases at no extra cost to you, which does not influence the independently determined ranking. As an Amazon Associate, BestLLMfor earns from qualifying purchases.
First open Apache 2.0 dLLM: MoE 16B/1B + 6.2B diffusion decoder. Unified text+vision. Released April 22, 2026.
# HuggingFace : inclusionAI/LLaDA2.0-Uni (Flash Attn 2 + CUDA 12.4 requis)
24B coding specialist, Apache 2.0. 72.2% SWE-Bench. 256k ctx, FR lab.
ollama run devstral-small2:24b
MoE variant of Gemma 4. 26B/4B active. Full multimodal (text+image+audio).
ollama run gemma4:26b
Dense multimodal 27B released April 22, 2026. 262k ctx (1M YaRN). SWE-bench Verified 77.2%.
ollama run qwen3.6:27b
Qwen 3.8 27B: dense multimodal (text + vision), 262k context, ~16 GB Q4 VRAM (18 GB of Ollama weights). Apache 2.0, agentic coding and vision.
ollama run qwen3.8:27b
First open Mistral reasoner. AIME24 70.7%. Based on Small 3.1 + CoT training.
ollama run magistral:24b
Dense 31B multimodal (text+image+audio). 140+ languages, 256k context. #3 open model on Chatbot Arena.
ollama run gemma4:31b
| Rank | Model | Params | Q4 VRAM | Context | License | On Radeon RX 7900 XT |
|---|---|---|---|---|---|---|
| #1 | LLaDA 2.0 Uni 16B | 16B | 18 GB | 8 192 | Apache 2.0 | 60 tok/s · Q4_K_M |
| #2 | Devstral Small 2 24B | 24B | 14 GB | 256 000 | Apache 2.0 | 15 tok/s · Q5_K_M |
| #3 | Gemma 4 26B-A4B MoE | 26B | 16 GB | 128 000 | Apache 2.0 | 22 tok/s · Q5_K_M |
| #4 | Qwen 3.6 27B | 27B | 16 GB | 262 144 | Apache 2.0 | 13 tok/s · Q5_K_M |
| #5 | Qwen 3.8 27B | 27B | 16 GB | 262 144 | Apache 2.0 | 14 tok/s · Q5_K_M |
| #6 | Magistral Small 24B | 24B | 14 GB | 128 000 | Apache 2.0 | 15 tok/s · Q5_K_M |
| #7 | Gemma 4 31B | 31B | 18 GB | 256 000 | Apache 2.0 | 12 tok/s · Q4_K_M |
Your private, free ChatGPT on your machine in 1 hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
Free memo
Get the memo VRAM → best coding model → Ollama command (one screen, copy and paste). Then switch to the Copilote Local kit for a setup that actually works.
The Local Copilot kit — the Ollama + Cline + Aider configs are ready to paste, with tuned Modelfiles, troubleshooting, and lifetime online access →No spam. Unsubscribe in 1 click. Your data stays with us (never resold).
Your card → the best coding model to run locally, and the exact Ollama command:
| Your VRAM | Typical GPUs / Macs | Recommended coding model | Command Ollama |
|---|---|---|---|
| 8 GB | RTX 4060 / 3060 · M1-M2 16 GB | Qwen 3.5 9B (Q4, 6.6 GB — 256k context) | ollama run qwen3.5:9b |
| 12 GB | RTX 3060 12 GB / 4070 / 5070 | Qwen 3.5 9B (Q8, 11 GB) or Gemma 4 12B (7.6 GB) | ollama run qwen3.5:9b-q8_0 |
| 16 GB | RTX 5070 Ti / 4080 / 5080 · RX 9070 XT · M4 24 GB | Devstral 24B (Q4, 14 GB) — coding-agent specialist | ollama run devstral:24b |
| 24 GB | RTX 3090 / 4090 · RX 7900 XTX · M4 Pro 48 GB | Qwen 3.8 27B (Q4, 18 GB) — the “close to Copilot” option | ollama run qwen3.8:27b |
| 32 GB | RTX 5090 | Qwen 3.6 35B-A3B (Q4, 23 GB) — fast MoE | ollama run qwen3.6:35b |
| 48 GB+ | Mac M4 Max 64 GB · M2 Ultra 128 GB | Qwen3-Coder 30B-A3B (Q8, 32 GB — 256k context) | ollama run qwen3-coder:30b-a3b-q8_0 |
-base : ollama run qwen2.5-coder:7b-base — it’s still the reference for this specific use case. ⚠️ Qwen 3.8: its reasoning is set very high by default and it “overthinks” simple requests — lower it to low (or turn it off) on first launch. ⚠️ License trap: Codestral 22B = Mistral Non-Production License → prohibited for coding at work. Qwen 3.5/3.8, Gemma 4, and Devstral are Apache 2.0. 💡 Running out of memory? Keep ~1.5 GB of VRAM free for context, or drop down one quantization level.🔌 To connect it to VS Code: Cline (multi-file agent), Aider (CLI) or Tabby/Twinny (FIM autocomplete) — they all connect to Ollama locally. The kit Local Copilot — ready-to-paste configs + tested setup — is available: /copilote-local.
Filter: Q4_K_M ≤ 18 GB. 13-32B bonus (20 GB peak). 800 GB/s + ROCm 6.
Criteria considered:
The scoring is fully transparent: see our methodology for details on VRAM/tokens/sec calculations.
RX 7900 XT versus 7900 XTX?
XT = 20 GB + 800 GB/s. XTX = 24 GB + 960 GB/s. For 24–32B Q4, both work. XTX is preferable for fine-tuning or 32B Q5. See RX 7900 XTX.
Why 20 GB instead of 16 or 24?
AMD’s marketing choice to position it between the 7900 XTX and the previous tier. For LLMs, it is the sweet spot: Mistral Small 24B Q4 (~13 GB) + 32k context + cache = ~18 GB.
ROCm on a 7900 XT: easy?
Yes, since ROCm 6 (2024). Ollama has native support via -gpu rocm. llama.cpp does too. Simpler than it was 2 years ago. See guide.
Sweet spot LLM for the 7900 XT?
Mistral Small 24B Q5 (~17 GB) at 25 tok/s, Qwen 3 32B Q4 (~17 GB) at 22 tok/s. Excellent for code + chat.
Prices in euros (€) are French market prices including VAT, checked by QuelLLM. US prices differ: the Amazon buttons show the current US price.