🇺🇸 Gemma 4 E2B
Gemma 4 E2B: 2B active (5.1B total), ~3 GB VRAM Q4 (weights Ollama 4.3 GB in QAT, 7.2 GB by default). Text-and-image multimodal, 128k context, Apache 2.0.
ollama run gemma4:e2b
Ranking updated on 09/10/2026
On an 8 GB Mac (M1/M2/M3 Air, base MacBook Air M4, base Mac mini M2), macOS takes ~4 GB. That leaves ~3–4 GB usable for an LLM. We limit ourselves to 1–3B models in Q4_K_M to keep things smooth.
Compare prices for Mac mini M5 Pro (24 GB / 512 GB) from our partner retailers (verified product pages):
Why this choice? Our complete guide to the Mac mini M5 Pro (24 GB / 512 GB) →
Which PC should you choose for your budget? Our picks from €800 to €3,500 →
Affiliate links — BestLLMfor may earn a commission on purchases, at no extra cost to you, which does not influence the ranking (established independently). As an Amazon Associate, BestLLMfor earns from qualifying purchases.
Gemma 4 E2B: 2B active (5.1B total), ~3 GB VRAM Q4 (weights Ollama 4.3 GB in QAT, 7.2 GB by default). Text-and-image multimodal, 128k context, Apache 2.0.
ollama run gemma4:e2b
Dense 3B Apache 2.0, 12 languages including FR, 131k ctx, GQA 40Q/8KV. Tool calling and code FIM. Released April 29, 2026.
ollama run granite4.1:3b
Granite 4.1 (3B Apache 2.0): generic Ollama tag from the IBM Granite 4.1 family, 128k ctx, tool calling, and code. Released May 2026.
ollama run granite4.1
3B VLM specialized in enterprise document extraction. OCR, tables, forms.
# HuggingFace : ibm-granite/granite-4.0-3b-vision
3B dual-mode (think/no-think). 6 languages. MMLU 59.7, GSM8K 70.9. Fully open (data + recipe).
# HuggingFace : HuggingFaceTB/SmolLM3-3B
MIT 3B OCR specialist. Praised 'optical compression' approach. DeepEncoder-based.
ollama run deepseek-ocr:3b
1.1B Apache 2.0 OpenBMB. Bilingual EN/ZH SFT with tool calling, optimized for on-device use. Q4 VRAM <1 GB for smartphones and modest laptops.
# HuggingFace : openbmb/MiniCPM5-1B-SFT
| Rank | Model | Params | Q4 VRAM | Context | License | On Apple M2 (16 GB) |
|---|---|---|---|---|---|---|
| #1 | Gemma 4 E2B | 2B | 3 GB | 128 000 | Apache 2.0 | 20 tok/s · FP16 |
| #2 | Granite 4.1 3B Instruct | 3B | 2 GB | 131 072 | Apache 2.0 | 25 tok/s · FP16 |
| #3 | Granite 4.1 | 3B | 1.7 GB | 128 000 | Apache 2.0 | 50 tok/s · FP16 |
| #4 | Granite 4.0 3B Vision | 3B | 2.2 GB | 16 384 | Apache 2.0 | 25 tok/s · FP16 |
| #5 | SmolLM3 3B | 3B | 2 GB | 128 000 | Apache 2.0 | 25 tok/s · FP16 |
| #6 | DeepSeek-OCR | 3B | 2 GB | 8 192 | MIT | 25 tok/s · FP16 |
| #7 | MiniCPM5 1B SFT | 1.1B | 0.6 GB | 32 768 | Apache 2.0 | 45 tok/s · FP16 |
Local AI on your Mac, fully explored: unified memory, MLX vs. GGUF, the right model for your chip, Ollama and LM Studio tuned for Apple Silicon.
Free memo
Get the memo VRAM → best coding model → Ollama command (one screen, copy and paste). Then switch to the Copilote Local kit for a setup that actually works.
The Local Copilot kit — the Ollama + Cline + Aider configs are ready to paste, with tuned Modelfiles, troubleshooting, and lifetime online access →No spam. Unsubscribe in 1 click. Your data stays with us (never resold).
Your card → the best coding model to run locally, and the exact Ollama command:
| Your VRAM | Typical GPUs / Macs | Recommended coding model | Command Ollama |
|---|---|---|---|
| 8 GB | RTX 4060 / 3060 · M1-M2 16 GB | Qwen 3.5 9B (Q4, 6.6 GB — 256k context) | ollama run qwen3.5:9b |
| 12 GB | RTX 3060 12 GB / 4070 / 5070 | Qwen 3.5 9B (Q8, 11 GB) or Gemma 4 12B (7.6 GB) | ollama run qwen3.5:9b-q8_0 |
| 16 GB | RTX 5070 Ti / 4080 / 5080 · RX 9070 XT · M4 24 GB | Devstral 24B (Q4, 14 GB) — coding-agent specialist | ollama run devstral:24b |
| 24 GB | RTX 3090 / 4090 · RX 7900 XTX · M4 Pro 48 GB | Qwen 3.8 27B (Q4, 18 GB) — the “close to Copilot” option | ollama run qwen3.8:27b |
| 32 GB | RTX 5090 | Qwen 3.6 35B-A3B (Q4, 23 GB) — fast MoE | ollama run qwen3.6:35b |
| 48 GB+ | Mac M4 Max 64 GB · M2 Ultra 128 GB | Qwen3-Coder 30B-A3B (Q8, 32 GB — 256k context) | ollama run qwen3-coder:30b-a3b-q8_0 |
-base : ollama run qwen2.5-coder:7b-base — it’s still the reference for this specific use case. ⚠️ Qwen 3.8: its reasoning is set very high by default and it “overthinks” simple requests — lower it to low (or turn it off) on first launch. ⚠️ License trap: Codestral 22B = Mistral Non-Production License → prohibited for coding at work. Qwen 3.5/3.8, Gemma 4, and Devstral are Apache 2.0. 💡 Running out of memory? Keep ~1.5 GB of VRAM free for context, or drop down one quantization level.🔌 To connect it to VS Code: Cline (multi-file agent), Aider (CLI) or Tabby/Twinny (FIM autocomplete) — they all connect to Ollama locally. The kit Local Copilot — ready-to-paste configs + tested setup — is available: /copilote-local.
Filter: 1-4B models whose Q4_K_M fits under 4 GB (leaving 4 GB for macOS + context). Bonus: 1-3B (peak 8 GB) and ≤ 2B (zero swap). Phi-4 Mini, Llama 3.2 3B, Gemma 4 3B dominate.
Criteria considered:
The scoring is fully transparent: see our methodology for details on VRAM/tokens/sec calculations.
Can an 8 GB Mac really run an LLM?
Yes, but tightly. Phi-4 Mini 3.8B Q4 (~2.3 GB) or Llama 3.2 3B Q4 (~2 GB) run at 25-35 tokens/sec. macOS takes 4 GB, leaving you 2-3 GB free — tight but usable for short chats.
Mac mini M2 8 GB vs. MacBook Air M4 16 GB?
The Air M4 16 GB is clearly preferable: 2× the RAM supports much more capable 7–8B models (Mistral, Qwen 3). The mini M2 8 GB does not comfortably exceed 3B. See 16 GB Mac.
Which quantization on 8 GB?
Q4_K_M remains the sweet spot. Q3_K_M can fit a Mistral 7B (~3.5 GB), but quality drops noticeably. Prefer a well-supported 3B Q4 over a crippled 7B Q3.
Do you really need 16 GB to get started with a local LLM?
For serious work, yes — Apple actually banned 8 GB on all M4 Macs in 2025. For occasional testing on an existing Mac, 8 GB is enough to explore 1–3B models. See MacBook Air M1.
Prices in euros (€) are French market prices including VAT, as checked by BestLLMfor. US prices differ: the Amazon buttons show the current US price.