🇺🇸 Gemma 4 26B-A4B MoE
MoE variant of Gemma 4. 26B/4B active. Full multimodal (text+image+audio).
ollama run gemma4:26b
Ranking updated on 09/10/2026
The Apple Silicon architecture (M1 to M4) shares memory between the CPU and GPU—excellent for LLMs. 7-32B models run remarkably well on Mac, especially Pro/Max models with 32-128 GB of unified memory.
Compare prices for Mac mini M5 Pro (24 GB / 512 GB) from our partner retailers (verified product pages):
Why this choice? Our complete guide to the Mac mini M5 Pro (24 GB / 512 GB) →
Which PC should you choose for your budget? Our picks from €800 to €3,500 →
Affiliate links — QuelLLM may earn a commission on purchases at no extra cost to you, which does not influence the ranking (established independently). As an Amazon Associate, BestLLMfor earns from qualifying purchases.
MoE variant of Gemma 4. 26B/4B active. Full multimodal (text+image+audio).
ollama run gemma4:26b
First open Apache 2.0 dLLM: MoE 16B/1B + 6.2B diffusion decoder. Unified text+vision. Released April 22, 2026.
# HuggingFace : inclusionAI/LLaDA2.0-Uni (Flash Attn 2 + CUDA 12.4 requis)
Dense multimodal 27B released April 22, 2026. 262k ctx (1M YaRN). SWE-bench Verified 77.2%.
ollama run qwen3.6:27b
Qwen 3.8 27B: dense multimodal (text + vision), 262k context, ~16 GB Q4 VRAM (18 GB of Ollama weights). Apache 2.0, agentic coding and vision.
ollama run qwen3.8:27b
GLM-4.7-Flash (MoE 31B, ~3B active): the best code/VRAM ratio in the 30B class. MIT, 128k ctx, very fast on 3090/4090.
ollama run glm-4.7-flash
Dense 31B multimodal (text+image+audio). 140+ languages, 256k context. #3 open model on Chatbot Arena.
ollama run gemma4:31b
Dense 30B Apache 2.0, 12 languages including FR, 131k ctx, GQA 32Q/8KV. OpenAI-compatible tool calling. Released April 29, 2026.
ollama run granite4.1:30b
DiffusionGemma 26B (Google): Gemma diffusion-based vision-language model, instruct, 128k context, 15 GB VRAM Q4. Apache 2.0. Released June 2026.
# HuggingFace : google/diffusiongemma-26B-A4B-it
| Rank | Model | Params | Q4 VRAM | Context | License | On Apple M4 Pro (48 GB) |
|---|---|---|---|---|---|---|
| #1 | Gemma 4 26B-A4B MoE | 26B | 16 GB | 128 000 | Apache 2.0 | 22 tok/s · Q8 |
| #2 | LLaDA 2.0 Uni 16B | 16B | 18 GB | 8 192 | Apache 2.0 | 60 tok/s · Q8 |
| #3 | Qwen 3.6 27B | 27B | 16 GB | 262 144 | Apache 2.0 | 13 tok/s · Q8 |
| #4 | Qwen 3.8 27B | 27B | 16 GB | 262 144 | Apache 2.0 | 14 tok/s · Q8 |
| #5 | GLM 4.7 Flash | 31B | 19 GB | 128 000 | MIT | 40 tok/s · Q8 |
| #6 | Gemma 4 31B | 31B | 18 GB | 256 000 | Apache 2.0 | 12 tok/s · Q8 |
| #7 | Granite 4.1 30B Instruct | 30B | 17 GB | 131 072 | Apache 2.0 | 12 tok/s · Q8 |
| #8 | DiffusionGemma 26B-A4B Instruct | 26B | 15 GB | 128 000 | Apache 2.0 | 14 tok/s · Q8 |
You found the best models for Apple Silicon. The Mac kit provides the complete table by chip and memory capacity (ch. 4) and explains when MLX really makes a difference (ch. 3).
Free memo
Get the memo VRAM → best coding model → Ollama command (one screen, copy and paste). Then switch to the Copilote Local kit for a setup that actually works.
The Local Copilot kit — the Ollama + Cline + Aider configs are ready to paste, with tuned Modelfiles, troubleshooting, and lifetime online access →No spam. Unsubscribe in 1 click. Your data stays with us (never resold).
Your card → the best coding model to run locally, and the exact Ollama command:
| Your VRAM | Typical GPUs / Macs | Recommended coding model | Command Ollama |
|---|---|---|---|
| 8 GB | RTX 4060 / 3060 · M1-M2 16 GB | Qwen 3.5 9B (Q4, 6.6 GB — 256k context) | ollama run qwen3.5:9b |
| 12 GB | RTX 3060 12 GB / 4070 / 5070 | Qwen 3.5 9B (Q8, 11 GB) or Gemma 4 12B (7.6 GB) | ollama run qwen3.5:9b-q8_0 |
| 16 GB | RTX 5070 Ti / 4080 / 5080 · RX 9070 XT · M4 24 GB | Devstral 24B (Q4, 14 GB) — coding-agent specialist | ollama run devstral:24b |
| 24 GB | RTX 3090 / 4090 · RX 7900 XTX · M4 Pro 48 GB | Qwen 3.8 27B (Q4, 18 GB) — the “close to Copilot” option | ollama run qwen3.8:27b |
| 32 GB | RTX 5090 | Qwen 3.6 35B-A3B (Q4, 23 GB) — fast MoE | ollama run qwen3.6:35b |
| 48 GB+ | Mac M4 Max 64 GB · M2 Ultra 128 GB | Qwen3-Coder 30B-A3B (Q8, 32 GB — 256k context) | ollama run qwen3-coder:30b-a3b-q8_0 |
-base : ollama run qwen2.5-coder:7b-base — it’s still the reference for this specific use case. ⚠️ Qwen 3.8: its reasoning is set very high by default and it “overthinks” simple requests — lower it to low (or turn it off) on first launch. ⚠️ License trap: Codestral 22B = Mistral Non-Production License → prohibited for coding at work. Qwen 3.5/3.8, Gemma 4, and Devstral are Apache 2.0. 💡 Running out of memory? Keep ~1.5 GB of VRAM free for context, or drop down one quantization level.🔌 To connect it to VS Code: Cline (multi-file agent), Aider (CLI) or Tabby/Twinny (FIM autocomplete) — they all connect to Ollama locally. The kit Local Copilot — ready-to-paste configs + tested setup — is available: /copilote-local.
We exclude models < 3B (underutilized) and > 72B (they don’t fit on consumer Macs). Bonus points for 7–32B sizes — the sweet spot for MacBook Pro / Mac Studio — and open licenses (MLX often requires converting the weights).
Criteria considered:
The scoring is fully transparent: see our methodology for details on VRAM/tokens/sec calculations.
Ollama or MLX on Mac?
Ollama is the simplest (1 command). MLX is 20-30% faster but requires weight conversion and a bit of terminal work. LM Studio combines both (choose Ollama or MLX in the UI).
Which Mac should you use to run a 70B?
Mac Studio M2 Ultra (192 GB), M3 Max 128 GB, or M4 Max 128 GB. A 70B in Q4 = 40 GB + context, so 64 GB minimum is recommended. M4 Pro 48 GB can handle it in Q3 with tradeoffs.
Can a MacBook Air M2 with 16 GB run an LLM?
Yes — Mistral 7B Q4 (4–5 GB) or Gemma 2 9B Q4 (6 GB) run on an M2 with 16 GB. Expect 10–15 tokens/sec. See the dedicated guide.
MLX faster than llama.cpp on Mac?
Yes, generally 15–30% faster because MLX is native to Apple Silicon. But llama.cpp supports more models and quantizations. For everyday use: Ollama (llama.cpp). For maximum performance: MLX.