🇺🇸 DiffusionGemma 26B-A4B Instruct
DiffusionGemma 26B (Google): Gemma diffusion-based vision-language model, instruct, 128k context, 15 GB VRAM Q4. Apache 2.0. Released June 2026.
# HuggingFace : google/diffusiongemma-26B-A4B-it
Ranking updated on 09/10/2026
16 GB of VRAM is the ideal tier for quantized 13-24B LLMs. Target cards: RTX 4080/4080 Super, RTX 5080, 4070 Ti Super, 4060 Ti 16 GB, RX 7800 XT. Here are the best models for this range.
RTX 4080 : purchasing alternative available for local AI — RTX 5080 16 GB :
A mini PC is a complete machine: check the required memory and software compatibility. It does not replace macOS/MLX or CUDA.
Which PC should you choose for your budget? Our picks from €800 to €3,500 →
Affiliate links — QuelLLM may earn a commission on purchases at no extra cost to you, which does not influence the ranking (established independently). As an Amazon Associate, BestLLMfor earns from qualifying purchases.
DiffusionGemma 26B (Google): Gemma diffusion-based vision-language model, instruct, 128k context, 15 GB VRAM Q4. Apache 2.0. Released June 2026.
# HuggingFace : google/diffusiongemma-26B-A4B-it
4B effective multimodal (text+image+audio). 140 languages. For laptops and edge devices.
ollama run gemma4:e4b
24B coding specialist, Apache 2.0. 72.2% SWE-Bench. 256k ctx, FR lab.
ollama run devstral-small2:24b
First open Mistral reasoner. AIME24 70.7%. Based on Small 3.1 + CoT training.
ollama run magistral:24b
Little brother of gpt-oss 120B. 21B/3.6B active. Matches o3-mini on a laptop.
ollama run openai/gpt-oss:20b
Compact MoE reasoner, 21B/3B active. Apache 2.0. Fast thanks to 3B active.
ollama pull hf.co/baidu/ernie-4.5-21b-GGUF
MoE 26B/3B active parameters from a US lab. Fast thanks to the 3B active parameters. Apache 2.0.
ollama pull hf.co/arcee-ai/Trinity-Mini-26B-GGUF
| Rank | Model | Params | Q4 VRAM | Context | License | On RTX 4080 |
|---|---|---|---|---|---|---|
| #1 | DiffusionGemma 26B-A4B Instruct | 26B | 15 GB | 128 000 | Apache 2.0 | 14 tok/s · Q4_K_M |
| #2 | Gemma 4 E4B | 4B | 10 GB | 128 000 | Apache 2.0 | 40 tok/s · FP16 |
| #3 | Devstral Small 2 24B | 24B | 14 GB | 256 000 | Apache 2.0 | 15 tok/s · Q4_K_M |
| #4 | Magistral Small 24B | 24B | 14 GB | 128 000 | Apache 2.0 | 15 tok/s · Q4_K_M |
| #5 | gpt-oss 20B | 21B | 13 GB | 128 000 | Apache 2.0 | 55 tok/s · Q5_K_M |
| #6 | ERNIE 4.5 21B-A3B Thinking | 21B | 13 GB | 131 072 | Apache 2.0 | 40 tok/s · Q5_K_M |
| #7 | Trinity Mini 26B-A3B | 26B | 15 GB | 131 072 | Apache 2.0 | 40 tok/s · Q4_K_M |
Your private, free ChatGPT on your machine in 1 hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
Free memo
Get the memo VRAM → best coding model → Ollama command (one screen, copy and paste). Then switch to the Copilote Local kit for a setup that actually works.
The Local Copilot kit — the Ollama + Cline + Aider configs are ready to paste, with tuned Modelfiles, troubleshooting, and lifetime online access →No spam. Unsubscribe in 1 click. Your data stays with us (never resold).
Your card → the best coding model to run locally, and the exact Ollama command:
| Your VRAM | Typical GPUs / Macs | Recommended coding model | Command Ollama |
|---|---|---|---|
| 8 GB | RTX 4060 / 3060 · M1-M2 16 GB | Qwen 3.5 9B (Q4, 6.6 GB — 256k context) | ollama run qwen3.5:9b |
| 12 GB | RTX 3060 12 GB / 4070 / 5070 | Qwen 3.5 9B (Q8, 11 GB) or Gemma 4 12B (7.6 GB) | ollama run qwen3.5:9b-q8_0 |
| 16 GB | RTX 5070 Ti / 4080 / 5080 · RX 9070 XT · M4 24 GB | Devstral 24B (Q4, 14 GB) — coding-agent specialist | ollama run devstral:24b |
| 24 GB | RTX 3090 / 4090 · RX 7900 XTX · M4 Pro 48 GB | Qwen 3.8 27B (Q4, 18 GB) — the “close to Copilot” option | ollama run qwen3.8:27b |
| 32 GB | RTX 5090 | Qwen 3.6 35B-A3B (Q4, 23 GB) — fast MoE | ollama run qwen3.6:35b |
| 48 GB+ | Mac M4 Max 64 GB · M2 Ultra 128 GB | Qwen3-Coder 30B-A3B (Q8, 32 GB — 256k context) | ollama run qwen3-coder:30b-a3b-q8_0 |
-base : ollama run qwen2.5-coder:7b-base — it’s still the reference for this specific use case. ⚠️ Qwen 3.8: its reasoning is set very high by default and it “overthinks” simple requests — lower it to low (or turn it off) on first launch. ⚠️ License trap: Codestral 22B = Mistral Non-Production License → prohibited for coding at work. Qwen 3.5/3.8, Gemma 4, and Devstral are Apache 2.0. 💡 Running out of memory? Keep ~1.5 GB of VRAM free for context, or drop down one quantization level.🔌 To connect it to VS Code: Cline (multi-file agent), Aider (CLI) or Tabby/Twinny (FIM autocomplete) — they all connect to Ollama locally. The kit Local Copilot — ready-to-paste configs + tested setup — is available: /copilote-local.
We keep models that fit in Q4_K_M within 16 GB, favoring those that use VRAM efficiently (50-95%) — a sign that the hardware is being utilized.
Criteria considered:
The scoring is fully transparent: see our methodology for details on VRAM/tokens/sec calculations.
Which LLM on RTX 4080 16 GB?
Our top pick: DiffusionGemma 26B-A4B Instruct. For a good quality/throughput compromise, stick with 13–24B models quantized in Q4_K_M or Q5.
Can you use Q5 or Q8 on 16 GB?
Yes for an 8-14B model (Q5 of a 14B = ~10 GB, Q8 of an 8B = ~10 GB). Not for a 24B model (Q5 ≈ 17 GB, over budget). Q4 remains the option for 24B models.
Can Gemma 2 27B fit in 16 GB?
Only in Q4_K_M (≈ 16 GB) — right at the limit. It exceeds the limit in Q5 (20 GB). Prefer Mistral Small 3.1 24B in Q4 (14 GB) to leave some headroom.
RTX 4080 vs RTX 4070 Ti Great for LLMs?
Both have 16 GB, but the 4080 is 30–40% faster (tier 4 vs. 3 in our scoring). If your budget allows, the 4080 Super or 5080 is clearly better.