🇺🇸 Gemma 4 E4B
4B effective multimodal (text+image+audio). 140 languages. For laptops and edge devices.
ollama run gemma4:e4b
Ranking updated on 09/10/2026
16 GB of unified memory is the practical minimum for local AI on a Mac. macOS uses 4 GB, leaving ~10-11 GB for a Q4_K_M LLM. The 7-9B models (Mistral, Qwen 3, Gemma 4) are the sweet spot.
Compare prices for Mac mini M5 Pro (24 GB / 512 GB) from our partner retailers (verified product pages):
Why this choice? Our complete guide to the Mac mini M5 Pro (24 GB / 512 GB) →
Which PC should you choose for your budget? Our picks from €800 to €3,500 →
Affiliate links — BestLLMfor may earn a commission on purchases, at no extra cost to you, which does not influence the ranking (established independently). As an Amazon Associate, BestLLMfor earns from qualifying purchases.
4B effective multimodal (text+image+audio). 140 languages. For laptops and edge devices.
ollama run gemma4:e4b
Dense 8B Apache 2.0, 12 languages including FR, 131k context, GQA 32Q/8KV. MMLU 73.84, HumanEval 85.37. Released April 29, 2026.
# HuggingFace : ibm-granite/granite-4.1-8b
Granite 4.2 8B (IBM): dense Apache 2.0, 128k context, ~4.6 GB Q4 VRAM. Multilingual chat, coding, and reasoning for the enterprise.
ollama pull granite4.2
SFT “thinking” fine-tune of OLMo 3 7B: step-by-step reasoning, 16k context, ~4.2 GB VRAM in Q4. 100% open, Apache 2.0 license.
# HuggingFace : zimplex/olmo3-7b-think-sft-eosfix-16k-3ep-euc
GLM 5.3 (Zhipu): dense 7B specialized in code and reasoning, 128k context, ~4.1 GB VRAM in Q4. Lightweight, runs on a 6–8 GB GPU, MIT license.
ollama pull glm-5.3
Dense 7B 100% open (weights + data + code). Complete transparency for research.
ollama run olmo-3:7b
Hybrid thinking/fast mode. 119 languages, 32k native (131k via YaRN).
ollama run qwen3:8b
Next-generation dense 9B. 262k ctx, improved hybrid thinking.
ollama run qwen3.5:9b
| Rank | Model | Params | Q4 VRAM | Context | License | On Apple M2 (16 GB) |
|---|---|---|---|---|---|---|
| #1 | Gemma 4 E4B | 4B | 10 GB | 128 000 | Apache 2.0 | 14 tok/s · Q4_K_M |
| #2 | Granite 4.1 8B Instruct | 8B | 5 GB | 131 072 | Apache 2.0 | 12 tok/s · Q8 |
| #3 | Granite 4.2 8B | 8B | 4.6 GB | 128 000 | Apache 2.0 | 32 tok/s · Q8 |
| #4 | OLMo 3 7B Think (SFT) | 7B | 4.2 GB | 16 000 | Apache 2.0 | 32 tok/s · Q8 |
| #5 | GLM 5.3 7B | 7B | 4.1 GB | 128 000 | MIT | 32 tok/s · Q8 |
| #6 | OLMo 3 7B | 7B | 5 GB | 8 192 | Apache 2.0 | 12 tok/s · Q8 |
| #7 | Qwen 3 8B | 8B | 5 GB | 131 072 | Apache 2.0 | 12 tok/s · Q8 |
| #8 | Qwen 3.5 9B | 9B | 6 GB | 262 000 | Apache 2.0 | 9 tok/s · Q8 |
Local AI on your Mac, fully explored: unified memory, MLX vs. GGUF, the right model for your chip, Ollama and LM Studio tuned for Apple Silicon.
Free memo
Get the memo VRAM → best coding model → Ollama command (one screen, copy and paste). Then switch to the Copilote Local kit for a setup that actually works.
The Local Copilot kit — the Ollama + Cline + Aider configs are ready to paste, with tuned Modelfiles, troubleshooting, and lifetime online access →No spam. Unsubscribe in 1 click. Your data stays with us (never resold).
Your card → the best coding model to run locally, and the exact Ollama command:
| Your VRAM | Typical GPUs / Macs | Recommended coding model | Command Ollama |
|---|---|---|---|
| 8 GB | RTX 4060 / 3060 · M1-M2 16 GB | Qwen 3.5 9B (Q4, 6.6 GB — 256k context) | ollama run qwen3.5:9b |
| 12 GB | RTX 3060 12 GB / 4070 / 5070 | Qwen 3.5 9B (Q8, 11 GB) or Gemma 4 12B (7.6 GB) | ollama run qwen3.5:9b-q8_0 |
| 16 GB | RTX 5070 Ti / 4080 / 5080 · RX 9070 XT · M4 24 GB | Devstral 24B (Q4, 14 GB) — coding-agent specialist | ollama run devstral:24b |
| 24 GB | RTX 3090 / 4090 · RX 7900 XTX · M4 Pro 48 GB | Qwen 3.8 27B (Q4, 18 GB) — the “close to Copilot” option | ollama run qwen3.8:27b |
| 32 GB | RTX 5090 | Qwen 3.6 35B-A3B (Q4, 23 GB) — fast MoE | ollama run qwen3.6:35b |
| 48 GB+ | Mac M4 Max 64 GB · M2 Ultra 128 GB | Qwen3-Coder 30B-A3B (Q8, 32 GB — 256k context) | ollama run qwen3-coder:30b-a3b-q8_0 |
-base : ollama run qwen2.5-coder:7b-base — it’s still the reference for this specific use case. ⚠️ Qwen 3.8: its reasoning is set very high by default and it “overthinks” simple requests — lower it to low (or turn it off) on first launch. ⚠️ License trap: Codestral 22B = Mistral Non-Production License → prohibited for coding at work. Qwen 3.5/3.8, Gemma 4, and Devstral are Apache 2.0. 💡 Running out of memory? Keep ~1.5 GB of VRAM free for context, or drop down one quantization level.🔌 To connect it to VS Code: Cline (multi-file agent), Aider (CLI) or Tabby/Twinny (FIM autocomplete) — they all connect to Ollama locally. The kit Local Copilot — ready-to-paste configs + tested setup — is available: /copilote-local.
Filter: 1-13B models whose Q4_K_M fits under 10 GB (leaving 6 GB for macOS + a large context). Bonus 3-9B (16 GB peak). Small active MoE (Qwen 3 30B-A3B) bonus at the upper limit.
Criteria considered:
The scoring is fully transparent: see our methodology for details on VRAM/tokens/sec calculations.
16 GB Mac: Mistral 7B or Qwen 3 8B?
Qwen 3 8B is slightly more capable (reasoning, code) and fits in Q4 (~5 GB). Mistral 7B is faster (~25-30 tok/s vs 22-28). For French, Mistral retains the advantage. Both are excellent on 16 GB.
Mac mini M4 16 GB as a 24/7 LLM server?
Yes, excellent. Ollama + Open WebUI, port 11434 behind a reverse proxy. Mistral 7B Q4 or Qwen 3 8B Q4 at 30+ tok/s. Idle power draw 10W, load 35W. See Mac mini M4.
Can Qwen 3 30B-A3B run on 16 GB?
Exactly: Q4_K_M requires ~17 GB for the full model, but MoE loads only ~3 GB of active parameters. With mmap + light swap, it works but is slower (15-20 tok/s). 24 GB or 32 GB is much more comfortable. See 32 GB Mac.
Mac 16 GB vs PC RTX 4060 16 GB for LLMs?
The RTX 4060 16 GB is ~2× faster on 7–9B models (actual GDDR6 VRAM versus 100–120 GB/s unified memory). The Mac wins on quiet operation, battery life, and ease of installation. See GPU comparison.
Prices in euros (€) are French market prices including VAT, as checked by BestLLMfor. US prices differ: the Amazon buttons show the current US price.