🇺🇸 Gemma 4 26B-A4B MoE
MoE variant of Gemma 4. 26B/4B active. Full multimodal (text+image+audio).
ollama run gemma4:26b
Ranking updated on 09/10/2026
The MacBook Pro M3 Pro / Max (18-128 GB, 300-400 GB/s) remains an excellent laptop for local AI in 2026. Comfortable with 30B models in Q4 / 70B in Q3.
Compare prices for MacBook Pro M5 Pro — 24 GB / 1 TB from our partner retailers (verified product pages):
Which PC should you choose for your budget? Our picks from €800 to €3,500 →
Affiliate links — BestLLMfor may earn a commission from purchases at no extra cost to you, which does not influence the independently determined ranking. As an Amazon Associate, BestLLMfor earns from qualifying purchases.
MoE variant of Gemma 4. 26B/4B active. Full multimodal (text+image+audio).
ollama run gemma4:26b
First open Apache 2.0 dLLM: MoE 16B/1B + 6.2B diffusion decoder. Unified text+vision. Released April 22, 2026.
# HuggingFace : inclusionAI/LLaDA2.0-Uni (Flash Attn 2 + CUDA 12.4 requis)
GLM-4.7-Flash (MoE 31B, ~3B active): the best code/VRAM ratio in the 30B class. MIT, 128k ctx, very fast on 3090/4090.
ollama run glm-4.7-flash
Mamba-2 + MoE 32B/9B hybrid. ~70% less RAM in long contexts. Apache 2.0.
ollama run granite4:small-h
MoE 30B/3B active hybrid thinking. MMLU 81.4, AIME24 80.4. 100+ languages.
ollama run qwen3:30b-a3b
MoE 33B/3B active parameters, Apache 2.0, specializing in agentic coding. 68.2% SWE-Bench Verified, 128k ctx. Runs on a 36 GB Mac. Released April 28, 2026.
ollama run laguna-xs.2
Dense multimodal 27B released April 22, 2026. 262k ctx (1M YaRN). SWE-bench Verified 77.2%.
ollama run qwen3.6:27b
Qwen 3.8 27B: dense multimodal (text + vision), 262k context, ~16 GB Q4 VRAM (18 GB of Ollama weights). Apache 2.0, agentic coding and vision.
ollama run qwen3.8:27b
| Rank | Model | Params | Q4 VRAM | Context | License | On Apple M3 Max (64 GB) |
|---|---|---|---|---|---|---|
| #1 | Gemma 4 26B-A4B MoE | 26B | 16 GB | 128 000 | Apache 2.0 | 22 tok/s · Q8 |
| #2 | LLaDA 2.0 Uni 16B | 16B | 18 GB | 8 192 | Apache 2.0 | 60 tok/s · FP16 |
| #3 | GLM 4.7 Flash | 31B | 19 GB | 128 000 | MIT | 40 tok/s · Q8 |
| #4 | Granite 4.0 H-Small 32B-A9B | 32B | 19 GB | 128 000 | Apache 2.0 | 30 tok/s · Q8 |
| #5 | Qwen 3 30B-A3B | 30B | 19 GB | 131 072 | Apache 2.0 | 40 tok/s · Q8 |
| #6 | Laguna XS.2 | 33B | 19 GB | 131 072 | Apache 2.0 | 40 tok/s · Q8 |
| #7 | Qwen 3.6 27B | 27B | 16 GB | 262 144 | Apache 2.0 | 13 tok/s · Q8 |
| #8 | Qwen 3.8 27B | 27B | 16 GB | 262 144 | Apache 2.0 | 14 tok/s · Q8 |
Local AI on your Mac, fully explored: unified memory, MLX vs. GGUF, the right model for your chip, Ollama and LM Studio tuned for Apple Silicon.
Free memo
Get the memo VRAM → best coding model → Ollama command (one screen, copy and paste). Then switch to the Copilote Local kit for a setup that actually works.
The Local Copilot kit — the Ollama + Cline + Aider configs are ready to paste, with tuned Modelfiles, troubleshooting, and lifetime online access →No spam. Unsubscribe in 1 click. Your data stays with us (never resold).
Your card → the best coding model to run locally, and the exact Ollama command:
| Your VRAM | Typical GPUs / Macs | Recommended coding model | Command Ollama |
|---|---|---|---|
| 8 GB | RTX 4060 / 3060 · M1-M2 16 GB | Qwen 3.5 9B (Q4, 6.6 GB — 256k context) | ollama run qwen3.5:9b |
| 12 GB | RTX 3060 12 GB / 4070 / 5070 | Qwen 3.5 9B (Q8, 11 GB) or Gemma 4 12B (7.6 GB) | ollama run qwen3.5:9b-q8_0 |
| 16 GB | RTX 5070 Ti / 4080 / 5080 · RX 9070 XT · M4 24 GB | Devstral 24B (Q4, 14 GB) — coding-agent specialist | ollama run devstral:24b |
| 24 GB | RTX 3090 / 4090 · RX 7900 XTX · M4 Pro 48 GB | Qwen 3.8 27B (Q4, 18 GB) — the “close to Copilot” option | ollama run qwen3.8:27b |
| 32 GB | RTX 5090 | Qwen 3.6 35B-A3B (Q4, 23 GB) — fast MoE | ollama run qwen3.6:35b |
| 48 GB+ | Mac M4 Max 64 GB · M2 Ultra 128 GB | Qwen3-Coder 30B-A3B (Q8, 32 GB — 256k context) | ollama run qwen3-coder:30b-a3b-q8_0 |
-base : ollama run qwen2.5-coder:7b-base — it’s still the reference for this specific use case. ⚠️ Qwen 3.8: its reasoning is set very high by default and it “overthinks” simple requests — lower it to low (or turn it off) on first launch. ⚠️ License trap: Codestral 22B = Mistral Non-Production License → prohibited for coding at work. Qwen 3.5/3.8, Gemma 4, and Devstral are Apache 2.0. 💡 Running out of memory? Keep ~1.5 GB of VRAM free for context, or drop down one quantization level.🔌 To connect it to VS Code: Cline (multi-file agent), Aider (CLI) or Tabby/Twinny (FIM autocomplete) — they all connect to Ollama locally. The kit Local Copilot — ready-to-paste configs + tested setup — is available: /copilote-local.
Filter: 3–100B with Q4_K_M fitting under 70 GB. Bonus: 13–70B (peak M3 Max) and 7–32B (M3 Pro). Well-rated MoE models.
Criteria considered:
The scoring is fully transparent: see our methodology for details on VRAM/tokens/sec calculations.
M3 Pro 18 GB MBP: enough for 13B?
Yes — Mistral Nemo 12B Q4 (~7 GB) at 28-35 tok/s. Mistral Small 24B Q4 (~13 GB) just manages 22-26 tok/s. See the MBP M3 guide.
Can an MBP M3 Max 128 GB run Llama 70B?
Yes—Llama 3.3 70B Q4_K_M (~40 GB) runs at 10–14 tok/s on an M3 Max. Q5_K_M (~48 GB) remains smooth. It was the first laptop that was practically capable.
M3 Max vs M4 Max?
M4 Max is ~15–20% faster with equivalent RAM (enhanced Neural Engine, 546 GB/s memory on 16 cores). M3 Max remains excellent: Llama 70B Q4 = 12 tok/s vs 15 tok/s on M4 Max. Not a necessary upgrade.
Which model codes on an M3 MacBook Pro?
Qwen 2.5 Coder 32B Q4 (~17 GB) on an M3 Pro 36 GB or M3 Max — excellent for Python/JS/Go coding. DeepSeek Coder V2 16B Q4 (~9 GB) is faster. See code ranking.
Prices in euros (€) are French market prices including VAT, checked by QuelLLM. US prices differ: the Amazon buttons show the current US price.