Intermediate 12 minMacBook Pro

Which LLM on a MacBook Pro M3 Pro / Max (18–128 GB) ?

The MacBook Pro M3 Pro / Max (late 2023) is the 2026 sweet spot for a local-LLM laptop. The 16-core M3 Max offers 400 GB/s of bandwidth and up to 128 GB of unified memory—enough to run the entire 2026 open-weight catalog at full precision: Qwen 3.6 35B-A3B in Q8, Qwen 3.8 27B in dense form, or multiple models loaded in parallel with 128k-plus contexts. Note: the M3 Pro suffered a memory downgrade (150 GB/s, versus 200 for the M2 Pro). This guide breaks down what each variant can really do.

Choosing a machine? Our picks by budget →

By Mohamed Meguedmi·Update 2026-08-27·Tested on macOS 14+
Recommended hardware

Buying alternative for this guide: GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395).

A mini PC is a complete machine: check the available memory and engine compatibility. It does not replace macOS/MLX or CUDA.

Why this choice? Our complete guide on GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395) →

Compare all options by budget, from €800 to €3,500 →

On the go: which laptop for local AI →

Affiliate links — possible commission at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.

#M3 Pro / Max MacBook Pro in 2026

Output
October–November 2023. Still supported by macOS Sequoia and later.
M3 Pro
11 or 12 CPU cores + 14 or 18 GPU cores, 150 GB/s (downgrade vs. M2 Pro), 18 or 36 GB RAM.
14-core M3 Max
14 CPU cores + 30 GPU cores, 300 GB/s, 36 or 96 GB of RAM.
M3 Max 16-core
16 CPU cores + 40 GPU cores, 400 GB/s, 48, 64, or 128 GB RAM.
!
The M3 Pro trap
Apple reduced the M3 Pro's memory bandwidth (150 GB/s versus 200 for M2 Pro). Result: the M3 Pro is slower than the M2 Pro for LLMs. If you want an M3, go straight to the M3 Max — or keep a used M2 Pro.

#1. M3 Pro vs. M3 Max (watch out for the trap)

M3 Pro 150 GB/s
Good for 9B–12B (~30 tok/s), but it runs out of steam on dense 24B+ models (~12 tok/s). Avoid it if you are targeting dense 24B+ models.
M3 Max 14-core 300 GB/s
The sweet spot: Qwen 3.6 35B-A3B (MoE) at 35–38 tok/s, Qwen 3.8 27B dense at 15–17 tok/s. Maximum RAM: 96 GB.
M3 Max 16-core 400 GB/s
The high-end option. +30% speed on large models, unlocking 128 GB of RAM for full Q8 precision and multiple models in parallel.

#2. 18 to 128 GB: which LLMs

Compatible models by M3 MBP configuration
ModelQuantM3 Pro 18M3 Pro 36M3 Max 64M3 Max 96M3 Max 128
Qwen 3.5 9BQ5_K_M✓ 30✓ 32✓ 46✓ 48✓ 50
Gemma 4 12BQ5_K_M✓ 14✓ 16✓ 26✓ 28✓ 30
Mistral Small 24BQ5_K_M—✓ 10✓ 17✓ 19✓ 20
Qwen 3.8 27BQ5_K_M——✓ 15✓ 16✓ 17
Qwen 3.6 35B-A3BQ5_K_M——✓ 38✓ 40✓ 42
Qwen 3.6 35B-A3BQ8_0——✓ 32✓ 34✓ 36
Qwen3-Coder 30B-A3BQ8_0——✓ 33✓ 35✓ 37
GLM 4.7 FlashQ4_K_M——✓ 44✓ 46✓ 48

#3. Tokens/sec benchmarks

Order of magnitude estimates — MBP M3 Max 16-core GPU, 128 GB, Ollama, power draw
ModelQ4_K_MQ5_K_MFlash Attn 8k
Qwen 3.5 9B54 t/s50 t/s48 t/s
Gemma 4 12B34 t/s30 t/s29 t/s
Mistral Small 24B22 t/s20 t/s19 t/s
Qwen 3.8 27B19 t/s17 t/s16 t/s
Qwen 3.6 35B-A3B46 t/s42 t/s40 t/s

#4. Installation and configuration

Complete MBP M3 Max 128 GB setup
brew install --cask ollama

# Relever la mémoire GPU (128 Go → 116 Go GPU)
sudo sysctl iogpu.wired_limit_mb=118784
echo "iogpu.wired_limit_mb=118784" | sudo tee -a /etc/sysctl.conf

# Flash Attention + KV cache Q8
export OLLAMA_FLASH_ATTENTION=1
export OLLAMA_KV_CACHE_TYPE=q8_0

brew services restart ollama
ollama pull qwen3.6:35b

#5. Large models at full precision with Flash Attention

Flash Attention, long limited to CUDA, has been available on Metal since late 2024 via llama.cpp and Ollama. Major gains on long contexts — essential for pushing a Qwen 3.6 35B-A3B to a 128k context.

The Mac Kit

Local AI on your Mac, done right: the models that fit your unified memory, MLX vs Ollama vs LM Studio, and the settings that make Apple Silicon fast.

  • Lifetime online access
  • PDF + files
  • Lifetime updates
Self-compiled llama.cpp — Qwen 3.6 35B 16k context
./build/bin/llama-cli \
  -m ./models/qwen3.6-35b-a3b-q5_K_M.gguf \
  -ngl 99 -fa \
  -ctk q8_0 -ctv q8_0 \
  -c 16384 \
  --temp 0.7
→
Measured gains with Flash Attention on M3 Max
4k context: +8%. 8k context: +15%. 16k context: +25%. The more context you have, the more noticeable the gain. Always enable this on M3 Max.

#6. The M3 Max 128 GB: workstation territory

Qwen 3.6 35B-A3B Q8
Full precision, maximum quality. ~36 tok/s (MoE, 3B active). The 128 GB enables Q8 with no compromise on context.
Qwen 3.8 27B Q8
The largest dense general-purpose model at full precision (~29 GB). ~15 tok/s. The “Copilot-like” model for 2026, with vision included.
2 30B models in parallel
One Qwen 3.6 35B-A3B + one Qwen3-Coder 30B-A3B loaded simultaneously. ~42 GB total. For multi-specialist agents.
128k+ context on 35B
With Q8 KV cache, Qwen 3.6 35B-A3B plus a 128K context fits in ~50 GB. Serious analysis of long documents.

#M3 Max vs. M4 Max

Bandwidth
M3 Max 16-core = 400 GB/s. M4 Max 16-core = 546 GB/s (+36%).
Speed Qwen 3.6 35B-A3B
M3 Max ~40 tok/s, M4 Max ~54 tok/s. +35% in real-world use.
Maximum RAM
Identical: 128 GB on both top configurations.
Verdict
M4 Max for a new purchase. A used M3 Max at -25% of the price is also an excellent deal.

#Frequently asked questions

Is the M3 Pro really slower than the M2 Pro for LLMs?+
Yes, for dense 24B+ models, because of the bandwidth (150 GB/s vs 200). On Qwen 3.5 9B, the gap is small (~5%). On Mistral Small 24B, the M2 Pro is 10–15% faster. Advice: if you're buying used and targeting a dense 24B+ model, prefer the M2 Pro or jump straight to the M3 Max.
Which M3 MBP can run the large 2026 models?+
M3 Max 14-core 64 GB for a Qwen 3.6 35B-A3B in Q8 (~35 tok/s, MoE), M3 Max 16-core 96–128 GB for loading multiple models or pushing the context to 128k. The M3 Pro 36 GB is limited to the dense 24B model.
What is the practical difference between the M3 Max 14-core and 16-core?+
+30% bandwidth (300 vs. 400 GB/s), +33% GPU cores, and a 128 GB RAM option. On a large 35B MoE: +40% speed. On a 9B: only +8%. If you're targeting large models, the 16-core is a genuine investment.
Is 128 GB of RAM really useful in 2026?+
Yes, if you want: Qwen 3.6 35B-A3B in Q8 (full precision, better quality than Q4), two 30B models in parallel, or a 128k context. For standard use, 64 GB is more than enough.
Does Flash Attention really make a difference on M3 Max?+
Yes, 15 to 25% more speed on 8k–16k contexts, with no loss of quality. Enable it by default with OLLAMA_FLASH_ATTENTION=1.
Battery life during a chat with a large model?+
About 1h40 in continuous chat on the 16" M3 Max (power consumption ~50 W). Comfortable for a session on the go. In batch mode on a large model, expect up to 1h — use wall power instead.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.