🇺🇸 Gemma 4 26B-A4B MoE
MoE variant of Gemma 4. 26B/4B active. Full multimodal (text+image+audio).
ollama run gemma4:26b
Ranking updated on 09/10/2026
24 GB of unified memory (high-end MacBook Air M2/M3/M4, base M4 Pro, high-end iMac M4) unlocks 13-14B models in Q4 and 30B-A3B MoE models. The sweet spot for high-quality local inference.
Compare prices for Mac mini M5 Pro (24 GB / 512 GB) from our partner retailers (verified product pages):
Why this choice? Our complete guide to the Mac mini M5 Pro (24 GB / 512 GB) →
Which PC should you choose for your budget? Our picks from €800 to €3,500 →
Affiliate links — QuelLLM may earn a commission on purchases at no extra cost to you, which does not influence the ranking (established independently). As an Amazon Associate, BestLLMfor earns from qualifying purchases.
MoE variant of Gemma 4. 26B/4B active. Full multimodal (text+image+audio).
ollama run gemma4:26b
Dense 8B Apache 2.0, 12 languages including FR, 131k context, GQA 32Q/8KV. MMLU 73.84, HumanEval 85.37. Released April 29, 2026.
# HuggingFace : ibm-granite/granite-4.1-8b
Gemma 4 12B (Google): dense multimodal model (text, vision, audio), 256k context, ~7 GB Q4 VRAM. Apache 2.0, multilingual.
# HuggingFace : google/gemma-4-12B
Granite 4.2 8B (IBM): dense Apache 2.0, 128k context, ~4.6 GB Q4 VRAM. Multilingual chat, coding, and reasoning for the enterprise.
ollama pull granite4.2
SFT “thinking” fine-tune of OLMo 3 7B: step-by-step reasoning, 16k context, ~4.2 GB VRAM in Q4. 100% open, Apache 2.0 license.
# HuggingFace : zimplex/olmo3-7b-think-sft-eosfix-16k-3ep-euc
GLM 5.3 (Zhipu): dense 7B specialized in code and reasoning, 128k context, ~4.1 GB VRAM in Q4. Lightweight, runs on a 6–8 GB GPU, MIT license.
ollama pull glm-5.3
Dense 14B with hybrid thinking. Equals Qwen 2.5 32B Based on STEM/code.
ollama run qwen3:14b
MIT 14B reasoner. Beats R1-Distill-Llama-70B on AIME/GPQA with 50× fewer parameters.
ollama run phi4-reasoning:14b
| Rank | Model | Params | Q4 VRAM | Context | License | On Apple M4 Pro (48 GB) |
|---|---|---|---|---|---|---|
| #1 | Gemma 4 26B-A4B MoE | 26B | 16 GB | 128 000 | Apache 2.0 | 22 tok/s · Q8 |
| #2 | Granite 4.1 8B Instruct | 8B | 5 GB | 131 072 | Apache 2.0 | 35 tok/s · FP16 |
| #3 | Gemma 4 12B | 12B | 7 GB | 262 144 | Apache 2.0 | 28 tok/s · FP16 |
| #4 | Granite 4.2 8B | 8B | 4.6 GB | 128 000 | Apache 2.0 | 50 tok/s · FP16 |
| #5 | OLMo 3 7B Think (SFT) | 7B | 4.2 GB | 16 000 | Apache 2.0 | 50 tok/s · FP16 |
| #6 | GLM 5.3 7B | 7B | 4.1 GB | 128 000 | MIT | 50 tok/s · FP16 |
| #7 | Qwen 3 14B | 14B | 9 GB | 131 072 | Apache 2.0 | 20 tok/s · FP16 |
| #8 | Phi-4 Reasoning 14B | 14B | 9 GB | 32 768 | MIT | 20 tok/s · FP16 |
Local AI on your Mac, fully explored: unified memory, MLX vs. GGUF, the right model for your chip, Ollama and LM Studio tuned for Apple Silicon.
Free memo
Get the memo VRAM → best coding model → Ollama command (one screen, copy and paste). Then switch to the Copilote Local kit for a setup that actually works.
The Local Copilot kit — the Ollama + Cline + Aider configs are ready to paste, with tuned Modelfiles, troubleshooting, and lifetime online access →No spam. Unsubscribe in 1 click. Your data stays with us (never resold).
Your card → the best coding model to run locally, and the exact Ollama command:
| Your VRAM | Typical GPUs / Macs | Recommended coding model | Command Ollama |
|---|---|---|---|
| 8 GB | RTX 4060 / 3060 · M1-M2 16 GB | Qwen 3.5 9B (Q4, 6.6 GB — 256k context) | ollama run qwen3.5:9b |
| 12 GB | RTX 3060 12 GB / 4070 / 5070 | Qwen 3.5 9B (Q8, 11 GB) or Gemma 4 12B (7.6 GB) | ollama run qwen3.5:9b-q8_0 |
| 16 GB | RTX 5070 Ti / 4080 / 5080 · RX 9070 XT · M4 24 GB | Devstral 24B (Q4, 14 GB) — coding-agent specialist | ollama run devstral:24b |
| 24 GB | RTX 3090 / 4090 · RX 7900 XTX · M4 Pro 48 GB | Qwen 3.8 27B (Q4, 18 GB) — the “close to Copilot” option | ollama run qwen3.8:27b |
| 32 GB | RTX 5090 | Qwen 3.6 35B-A3B (Q4, 23 GB) — fast MoE | ollama run qwen3.6:35b |
| 48 GB+ | Mac M4 Max 64 GB · M2 Ultra 128 GB | Qwen3-Coder 30B-A3B (Q8, 32 GB — 256k context) | ollama run qwen3-coder:30b-a3b-q8_0 |
-base : ollama run qwen2.5-coder:7b-base — it’s still the reference for this specific use case. ⚠️ Qwen 3.8: its reasoning is set very high by default and it “overthinks” simple requests — lower it to low (or turn it off) on first launch. ⚠️ License trap: Codestral 22B = Mistral Non-Production License → prohibited for coding at work. Qwen 3.5/3.8, Gemma 4, and Devstral are Apache 2.0. 💡 Running out of memory? Keep ~1.5 GB of VRAM free for context, or drop down one quantization level.🔌 To connect it to VS Code: Cline (multi-file agent), Aider (CLI) or Tabby/Twinny (FIM autocomplete) — they all connect to Ollama locally. The kit Local Copilot — ready-to-paste configs + tested setup — is available: /copilote-local.
Filter: 3–32B models whose Q4_K_M fits under 16 GB (leaving 8 GB for macOS + context). Bonus: 7–14B (24 GB dense peak) and 30B-A3B MoE (Apple sweet spot).
Criteria considered:
The scoring is fully transparent: see our methodology for details on VRAM/tokens/sec calculations.
Mac 24 GB: the 2026 LLM sweet spot?
Yes for 1-user models. Qwen 3 14B Q4 (~8 GB), Mistral Nemo 12B Q4 (~7 GB), Qwen 3 30B-A3B (MoE, ~17 GB) — all run at 25-40 tokens/sec. For sustained dense 13B+, 32 GB or more.
MacBook Air M4 24 GB vs. Mac mini M4 24 GB?
Exactly the same M4 chip + 120 GB/s. Difference: Air = fanless (throttles after ~10 minutes of sustained generation), mini = actively cooled and therefore stable 24/7. See MBA M4 or mini M4.
Which model codes on a 24 GB Mac?
Qwen 2.5 Coder 14B Q4 (~8 GB) or DeepSeek Coder V2 16B Q4 (~9 GB)—excellent for Python/JS/Go. Qwen 3 14B for general-purpose use. See code ranking.
Is 24 GB enough for an assistant + RAG?
Yes: Mistral Nemo 12B Q4 (~7 GB) + ChromaDB (1–2 GB) + 32k context (~3 GB) = ~12 GB used. Comfortable headroom. See the RAG guide.