🇺🇸 Gemma 4 26B-A4B MoE
MoE variant of Gemma 4. 26B/4B active. Full multimodal (text+image+audio).
ollama run gemma4:26b
Ranking updated on 09/10/2026
32 GB of unified memory is the mainstream quality tier. You can comfortably run Mistral Small 24B Q4, Qwen 3 30B Q4, or Qwen 3 30B-A3B MoE in Q8.
Compare prices for Mac mini M5 Pro (24 GB / 512 GB) from our partner retailers (verified product pages):
Why this choice? Our complete guide to the Mac mini M5 Pro (24 GB / 512 GB) →
Which PC should you choose for your budget? Our picks from €800 to €3,500 →
Affiliate links — QuelLLM may earn a commission on purchases at no extra cost to you, which does not influence the ranking (established independently). As an Amazon Associate, BestLLMfor earns from qualifying purchases.
MoE variant of Gemma 4. 26B/4B active. Full multimodal (text+image+audio).
ollama run gemma4:26b
First open Apache 2.0 dLLM: MoE 16B/1B + 6.2B diffusion decoder. Unified text+vision. Released April 22, 2026.
# HuggingFace : inclusionAI/LLaDA2.0-Uni (Flash Attn 2 + CUDA 12.4 requis)
MoE 30B/3B active hybrid thinking. MMLU 81.4, AIME24 80.4. 100+ languages.
ollama run qwen3:30b-a3b
MoE with 30B/3B active: thinking mode + instruct. Gold medalist at IMO 2025 and IOI 2025. Fast inference thanks to the 3B active parameters, with 30B-level reasoning capabilities. Released April 2026.
ollama run nemotron-cascade-2
MoE 30B (3.3B active parameters) specialized in agentic coding. Very fast locally, native 256k ctx, the benchmark for 16–24 GB via Ollama.
ollama run qwen3-coder:30b
Vision MoE with 30B/3B active. Vision sweet spot Qwen 3. 256k ctx.
ollama run qwen3-vl:30b
Korean agentic MoE with 30B/3B active parameters. Covers KR/EN/JP/ZH/TH/VI. Apache 2.0. MLA attention.
ollama pull hf.co/kakaoai/Kanana-2-30B-GGUF
Omni MoE 30B/3B active. Speech streaming. 119 ASR languages. Apache 2.0.
ollama run qwen3-omni:30b
| Rank | Model | Params | Q4 VRAM | Context | License | On Apple M4 Pro (48 GB) |
|---|---|---|---|---|---|---|
| #1 | Gemma 4 26B-A4B MoE | 26B | 16 GB | 128 000 | Apache 2.0 | 22 tok/s · Q8 |
| #2 | LLaDA 2.0 Uni 16B | 16B | 18 GB | 8 192 | Apache 2.0 | 60 tok/s · Q8 |
| #3 | Qwen 3 30B-A3B | 30B | 19 GB | 131 072 | Apache 2.0 | 40 tok/s · Q8 |
| #4 | Nemotron Cascade 2 30B-A3B | 30B | 17 GB | 128 000 | NVIDIA Open Model License | 30 tok/s · Q8 |
| #5 | Qwen3-Coder 30B-A3B | 30B | 19 GB | 262 144 | Apache 2.0 | 40 tok/s · Q8 |
| #6 | Qwen 3 VL 30B-A3B | 30B | 19 GB | 262 144 | Apache 2.0 | 40 tok/s · Q8 |
| #7 | Kanana 2 30B-A3B Thinking | 30B | 18 GB | 131 072 | Apache 2.0 | 40 tok/s · Q8 |
| #8 | Qwen 3 Omni 30B-A3B | 30B | 19 GB | 131 072 | Apache 2.0 | 40 tok/s · Q8 |
Local AI on your Mac, fully explored: unified memory, MLX vs. GGUF, the right model for your chip, Ollama and LM Studio tuned for Apple Silicon.
Free memo
Get the memo VRAM → best coding model → Ollama command (one screen, copy and paste). Then switch to the Copilote Local kit for a setup that actually works.
The Local Copilot kit — the Ollama + Cline + Aider configs are ready to paste, with tuned Modelfiles, troubleshooting, and lifetime online access →No spam. Unsubscribe in 1 click. Your data stays with us (never resold).
Your card → the best coding model to run locally, and the exact Ollama command:
| Your VRAM | Typical GPUs / Macs | Recommended coding model | Command Ollama |
|---|---|---|---|
| 8 GB | RTX 4060 / 3060 · M1-M2 16 GB | Qwen 3.5 9B (Q4, 6.6 GB — 256k context) | ollama run qwen3.5:9b |
| 12 GB | RTX 3060 12 GB / 4070 / 5070 | Qwen 3.5 9B (Q8, 11 GB) or Gemma 4 12B (7.6 GB) | ollama run qwen3.5:9b-q8_0 |
| 16 GB | RTX 5070 Ti / 4080 / 5080 · RX 9070 XT · M4 24 GB | Devstral 24B (Q4, 14 GB) — coding-agent specialist | ollama run devstral:24b |
| 24 GB | RTX 3090 / 4090 · RX 7900 XTX · M4 Pro 48 GB | Qwen 3.8 27B (Q4, 18 GB) — the “close to Copilot” option | ollama run qwen3.8:27b |
| 32 GB | RTX 5090 | Qwen 3.6 35B-A3B (Q4, 23 GB) — fast MoE | ollama run qwen3.6:35b |
| 48 GB+ | Mac M4 Max 64 GB · M2 Ultra 128 GB | Qwen3-Coder 30B-A3B (Q8, 32 GB — 256k context) | ollama run qwen3-coder:30b-a3b-q8_0 |
-base : ollama run qwen2.5-coder:7b-base — it’s still the reference for this specific use case. ⚠️ Qwen 3.8: its reasoning is set very high by default and it “overthinks” simple requests — lower it to low (or turn it off) on first launch. ⚠️ License trap: Codestral 22B = Mistral Non-Production License → prohibited for coding at work. Qwen 3.5/3.8, Gemma 4, and Devstral are Apache 2.0. 💡 Running out of memory? Keep ~1.5 GB of VRAM free for context, or drop down one quantization level.🔌 To connect it to VS Code: Cline (multi-file agent), Aider (CLI) or Tabby/Twinny (FIM autocomplete) — they all connect to Ollama locally. The kit Local Copilot — ready-to-paste configs + tested setup — is available: /copilote-local.
Filter: 3–35B models whose Q4_K_M fits under 22 GB (leaving 10 GB for macOS + long context). Bonus 13–30B (32 GB peak) and MoE (Qwen 3 30B-A3B in Q8 here).
Criteria considered:
The scoring is fully transparent: see our methodology for details on VRAM/tokens/sec calculations.
32 GB Mac: can it run Mistral Small 24B in Q5?
Yes: Q5_K_M ~17 GB. At 18–25 tokens/sec depending on the chip (M1 Max is the fastest, base M4 the most efficient). Excellent for French + general-purpose tasks.
Qwen 3 30B-A3B (MoE) in Q4 or Q8 on 32 GB?
Q8 (~32 GB) just fits—it uses everything. Q4 (~17 GB) is more comfortable and frees 15 GB for context and other apps. The Q4 vs. Q8 quality difference on MoE is marginal (<2% on benchmarks). Prefer Q4.
32 GB vs 48 GB: what kind of quality leap?
48 GB unlocks 32B dense models in Q5 and 70B models in Q3. 32 GB remains limited to 30B models in Q4 or 30B-A3B MoE in Q8. If you're buying new, 48 GB is better. See 48 GB Mac.
Which French model on a 32 GB Mac?
Mistral Small 3.1 24B Q4 (~13 GB) or Mistral Small 3.2 24B Q4—the best choices for French. Magistral Small 24B for reasoning. See FR ranking.
Learn more with our detailed head-to-head matchups of the finalists: