Home › Catalog › Best LLM for Mac with 8 GB of unified memory in 2026

Best LLM for Mac with 8 GB of unified memory in 2026

◆ Mac — Local AI on your Mac, done right — MLX, Ollama, LM Studio on Apple Silicon · $24 · or all kits $49 →

Ranking updated on 09/10/2026

On an 8 GB Mac (M1/M2/M3 Air, base MacBook Air M4, base Mac mini M2), macOS takes ~4 GB. That leaves ~3–4 GB usable for an LLM. We limit ourselves to 1–3B models in Q4_K_M to keep things smooth.

Offers and alternatives for local AI

Compare prices for Mac mini M5 Pro (24 GB / 512 GB) from our partner retailers (verified product pages):

Why this choice? Our complete guide to the Mac mini M5 Pro (24 GB / 512 GB) →

Which PC should you choose for your budget? Our picks from €800 to €3,500 →

Affiliate links — BestLLMfor may earn a commission on purchases, at no extra cost to you, which does not influence the ranking (established independently). As an Amazon Associate, BestLLMfor earns from qualifying purchases.

Ranking

1

🇺🇸 Gemma 4 E2B

Google · 2B parameters · Apache 2.0 · 128,000 tokens ctx

Gemma 4 E2B: 2B active (5.1B total), ~3 GB VRAM Q4 (weights Ollama 4.3 GB in QAT, 7.2 GB by default). Text-and-image multimodal, 128k context, Apache 2.0.

Why this ranking Gemma 4 E2B: 2B active (5.1B total), ~3 GB VRAM Q4 (weights Ollama 4.3 GB in QAT, 7.2 GB by default). Text-and-image multimodal, 128k context, Apache 2.0.
ollama run gemma4:e2b
On Apple M2 (16 GB)
FP16
10 GB · 20 tok/s
2

🇺🇸 Granite 4.1 3B Instruct

IBM · 3B parameters · Apache 2.0 · 131,072 tokens ctx

Dense 3B Apache 2.0, 12 languages including FR, 131k ctx, GQA 40Q/8KV. Tool calling and code FIM. Released April 29, 2026.

Why this ranking Dense 3B Apache 2.0, 12 languages including FR, 131k ctx, GQA 40Q/8KV. Tool calling and code FIM. Released April 29, 2026.
ollama run granite4.1:3b
On Apple M2 (16 GB)
FP16
6 GB · 25 tok/s
3

🇺🇸 Granite 4.1

IBM · 3B parameters · Apache 2.0 · 128,000 tokens ctx

Granite 4.1 (3B Apache 2.0): generic Ollama tag from the IBM Granite 4.1 family, 128k ctx, tool calling, and code. Released May 2026.

Why this ranking Granite 4.1 (3B Apache 2.0): generic Ollama tag from the IBM Granite 4.1 family, 128k ctx, tool calling, and code. Released May 2026.
ollama run granite4.1
On Apple M2 (16 GB)
FP16
6 GB · 50 tok/s
4

🇺🇸 Granite 4.0 3B Vision

IBM · 3B parameters · Apache 2.0 · 16,384-token context

3B VLM specialized in enterprise document extraction. OCR, tables, forms.

Why this ranking 3B VLM specialized in enterprise document extraction. OCR, tables, forms.
# HuggingFace : ibm-granite/granite-4.0-3b-vision
On Apple M2 (16 GB)
FP16
6.5 GB · 25 tok/s
5

🇫🇷 SmolLM3 3B

HuggingFace · 3B parameters · Apache 2.0 · 128,000 tokens ctx

3B dual-mode (think/no-think). 6 languages. MMLU 59.7, GSM8K 70.9. Fully open (data + recipe).

Why this ranking 3B dual-mode (think/no-think). 6 languages. MMLU 59.7, GSM8K 70.9. Fully open (data + recipe).
# HuggingFace : HuggingFaceTB/SmolLM3-3B
On Apple M2 (16 GB)
FP16
6 GB · 25 tok/s
6

🇨🇳 DeepSeek-OCR

DeepSeek · 3B parameters · MIT · 8,192-token context

MIT 3B OCR specialist. Praised 'optical compression' approach. DeepEncoder-based.

Why this ranking MIT 3B OCR specialist. Praised 'optical compression' approach. DeepEncoder-based.
ollama run deepseek-ocr:3b
On Apple M2 (16 GB)
FP16
6 GB · 25 tok/s
7

🇨🇳 MiniCPM5 1B SFT

OpenBMB · 1.1B parameters · Apache 2.0 · 32,768-token context

1.1B Apache 2.0 OpenBMB. Bilingual EN/ZH SFT with tool calling, optimized for on-device use. Q4 VRAM <1 GB for smartphones and modest laptops.

Why this ranking 1.1B Apache 2.0 OpenBMB. Bilingual EN/ZH SFT with tool calling, optimized for on-device use. Q4 VRAM <1 GB for smartphones and modest laptops.
# HuggingFace : openbmb/MiniCPM5-1B-SFT
On Apple M2 (16 GB)
FP16
2.2 GB · 45 tok/s

Comparison table

Rank Model Params Q4 VRAM Context License On Apple M2 (16 GB)
#1 Gemma 4 E2B 2B 3 GB 128 000 Apache 2.0 20 tok/s · FP16
#2 Granite 4.1 3B Instruct 3B 2 GB 131 072 Apache 2.0 25 tok/s · FP16
#3 Granite 4.1 3B 1.7 GB 128 000 Apache 2.0 50 tok/s · FP16
#4 Granite 4.0 3B Vision 3B 2.2 GB 16 384 Apache 2.0 25 tok/s · FP16
#5 SmolLM3 3B 3B 2 GB 128 000 Apache 2.0 25 tok/s · FP16
#6 DeepSeek-OCR 3B 2 GB 8 192 MIT 25 tok/s · FP16
#7 MiniCPM5 1B SFT 1.1B 0.6 GB 32 768 Apache 2.0 45 tok/s · FP16
The Mac kit

Local AI on your Mac, fully explored: unified memory, MLX vs. GGUF, the right model for your chip, Ollama and LM Studio tuned for Apple Silicon.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

Free memo

Which coding model should you run on YOUR machine?

Get the memo VRAM → best coding model → Ollama command (one screen, copy and paste). Then switch to the Copilote Local kit for a setup that actually works.

The Local Copilot kit — the Ollama + Cline + Aider configs are ready to paste, with tuned Modelfiles, troubleshooting, and lifetime online access →

No spam. Unsubscribe in 1 click. Your data stays with us (never resold).

Ranking methodology

Filter: 1-4B models whose Q4_K_M fits under 4 GB (leaving 4 GB for macOS + context). Bonus: 1-3B (peak 8 GB) and ≤ 2B (zero swap). Phi-4 Mini, Llama 3.2 3B, Gemma 4 3B dominate.

Criteria considered:

  • Q4_K_M ≤ 4 GB
  • Zero-swap macOS
  • Tokens/sec ≥ 20
  • Usable 2–4k context

The scoring is fully transparent: see our methodology for details on VRAM/tokens/sec calculations.

Frequently asked questions

Can an 8 GB Mac really run an LLM?

Yes, but tightly. Phi-4 Mini 3.8B Q4 (~2.3 GB) or Llama 3.2 3B Q4 (~2 GB) run at 25-35 tokens/sec. macOS takes 4 GB, leaving you 2-3 GB free — tight but usable for short chats.

Mac mini M2 8 GB vs. MacBook Air M4 16 GB?

The Air M4 16 GB is clearly preferable: 2× the RAM supports much more capable 7–8B models (Mistral, Qwen 3). The mini M2 8 GB does not comfortably exceed 3B. See 16 GB Mac.

Which quantization on 8 GB?

Q4_K_M remains the sweet spot. Q3_K_M can fit a Mistral 7B (~3.5 GB), but quality drops noticeably. Prefer a well-supported 3B Q4 over a crippled 7B Q3.

Do you really need 16 GB to get started with a local LLM?

For serious work, yes — Apple actually banned 8 GB on all M4 Macs in 2025. For occasional testing on an existing Mac, 8 GB is enough to explore 1–3B models. See MacBook Air M1.

Go further

BestLLMfor Kits The reference guide by use case
All kits for life — $49

Prices in euros (€) are French market prices including VAT, as checked by BestLLMfor. US prices differ: the Amazon buttons show the current US price.