Which LLM on MacBook Pro M2 Pro / Max (16–96 GB) ?
A MacBook Pro M2 Pro comfortably runs an 8–9B model with 16 GB and a 24–32B model with 32 GB; an M2 Max with 64 or 96 GB is the only variant of this generation that can run a 70B model in Q4 (about 40 GB) and two 20–30B models at the same time. Speed depends on bandwidth: 200 GB/s for the M2 Pro and 400 GB/s for the M2 Max, yielding nearly twice as many tokens per second for the same model.
Released in January 2023, the MacBook Pro M2 Pro / Max remains one of the most capable laptops for local AI: fans, 400 GB/s of bandwidth on the M2 Max, and up to 96 GB of unified memory. This page explains what each memory configuration can handle, what the published measurements show, why doubling bandwidth does not double speed, and when it is worth moving to an M4 Max.
Choosing a machine? Our picks by budget →
Buying alternative for this guide: GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395).
A mini PC is a complete machine: check the available memory and engine compatibility. It does not replace macOS/MLX or CUDA.
Why this choice? Our complete guide on GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395) →
Compare all options by budget, from €800 to €3,500 →
On the go: which laptop for local AI →
Affiliate links — possible commission at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.
#MacBook Pro M2 Pro / Max in 2026: the useful technical specifications
When it was unveiled, Apple announced 200 GB/s of bandwidth and up to 32 GB of unified memory for the M2 Pro, and 400 GB/s, up to 96 GB, and a GPU with up to 38 cores for the M2 Max. Apple presented those 96 GB as pushing the limits of graphics memory in a laptop. The memory is soldered to the processor, so you choose the configuration when you buy it and cannot change it later.
| Chip | Bandwidth | Possible memory | GPU |
|---|---|---|---|
| M2 Pro | 200 GB/s | 16 or 32 GB | 16 or 19 cores |
| M2 Max | 400 GB/s | 32, 64, or 96 GB | 30 or 38 cores |
Active cooling is the second advantage over a MacBook Air: the MacBook Pro can sustain long workloads without the chip having to slow down significantly. Apple also offers, on some models, a High Power mode that lets the fans spin faster; it’s worth trying for long generations, at the cost of additional noise.
#M2 Pro or M2 Max: what bandwidth really changes
Local AI on your Mac, done right: the models that fit your unified memory, MLX vs Ollama vs LM Studio, and the settings that make Apple Silicon fast.
- Lifetime online access
- PDF + files
- Lifetime updates
Text generation reads all the model’s active weights for every token. Maximum speed is therefore bandwidth divided by weight size. An M2 Max at 400 GB/s could theoretically be twice as fast as an M2 Pro at 200 GB/s. Measurements published in the llama.cpp repository show that this is only partly true.
| Chip (GPU cores) | Bandwidth | Q4_0 | Q8_0 | F16 |
|---|---|---|---|---|
| M2 Pro (19) | 200 GB/s | 38,86 t/s | 23,01 t/s | 13,06 t/s |
| M2 Max (30) | 400 GB/s | 60,99 t/s | 39,97 t/s | 24,16 t/s |
| M2 Max (38) | 400 GB/s | 65,95 t/s | 41,83 t/s | 24,65 t/s |
On a 7B model in Q4_0, the M2 Max is only 1.6 to 1.7 times faster than the M2 Pro. On the same model in F16, around 13 GB, the gap rises to nearly 1.9. The calculation explains why: a small model is read so quickly that other limits (compute, latency, overhead) take over, while a large model genuinely saturates memory. The practical consequence is counterintuitive: the larger the model, the more the M2 Max justifies its price. For a 7B, an M2 Pro is more than enough.
#From 16 to 96 GB: what fits and what remains borderline
On Apple Silicon, the GPU does not have access to all memory by default: according to discussions in the llama.cpp repository, Metal grants it between two-thirds and three-quarters. On 16 GB, allow about 10 to 12 GB for the model and its context; on 96 GB, about 64 to 72 GB. The limit can be raised through a system setting detailed in the dedicated guide, but exceeding it may freeze the machine.
| Model | Q4 weight | M2 Pro 16 GB | M2 Pro 32 GB | M2 Max 64 GB | M2 Max 96 GB |
|---|---|---|---|---|---|
| Qwen 3.5 9B | 6 GB | Comfortable | Yes | Yes | Yes |
| Gemma 4 12B | 7 GB | Just | Yes | Yes | Yes |
| Mistral Small 3.2 24B | 14 GB | No | Yes | Yes | Yes |
| Qwen 3.5 27B | 16 GB | No | Yes, moderate context | Yes | Yes |
| Qwen3-Coder 30B-A3B (MoE) | 19 GB | No | Yes, moderate context | Yes | Yes |
| Qwen 3.6 35B-A3B (MoE) | 21 GB | No | Just | Yes | Yes |
| Llama 3.3 70B | 40 GB | No | No | Yes, short context | Yes |
MoE models such as Qwen3-Coder 30B-A3B or Qwen 3.6 35B-A3B weigh 19 and 21 GB, but activate only about 3 billion parameters per token: at the same memory capacity, they generate noticeably faster than a 27B dense model. This is the most cost-effective option at 32 GB. At 96 GB, the M2 Max is less useful for loading a single very large model than for keeping two models resident at once, such as a chat model and a coding model.
#Published throughput: what we can and cannot claim
We don't have in-house measurements for this machine, and tokens per second depend on the model, quantization, context, and the version of llama.cpp or Ollama. The only measurements we cite are those published by users on a Llama 7B, which provide the scale. For a newer model, the bandwidth-divided-by-weights calculation remains the right rule of thumb, with the margin of error indicated above.
Prompt processing follows a different pattern: the llama.cpp table reports, for the same Llama 7B in F16, 384 t/s for prompt processing on a 19-core M2 Pro and 756 t/s on a 38-core M2 Max. The number of GPU cores affects this phase, not generation. For RAG or long-document analysis, an M2 Max will therefore wait about half as long for the first word as an M2 Pro.
#Install Ollama and use it correctly on a Mac
Ollama requires macOS Sonoma (14) or later. On Mac, the official FAQ specifies that when Ollama runs as an application, its environment variables must be set with launchctl, and then the application must be restarted. Exports in a terminal are not enough; the application does not see them.
The Ollama FAQ states that the q8_0-quantized K/V cache uses roughly half the memory of the default f16 format, and that Flash Attention must be enabled to benefit from it. On a 70B model with a long context, this setting is what makes the cache fit. One limitation to know: cache quantization can slightly reduce quality on some models; test it on your use cases before keeping it.
#Which model for which use case
- General-purpose chat
- An 8B to 12B model (Qwen 3.5 9B, Gemma 4 12B) responds quickly on an M2 Pro; a 27B model becomes worthwhile starting at 32 GB for more polished French.
- Code
- A coding MoE such as Qwen3-Coder 30B-A3B on 32 GB or more, or a 24B coding-agent model such as Devstral Small 2 (14 GB in Q4). For online autocompletion, a 7B model is sufficient.
- Document RAG
- Gemma 4 12B for speed, a 24B model for reasoning quality; limit the context to avoid extending the wait before the first token.
- Long analysis, agents
- On 64 GB or more: Qwen 3.6 35B-A3B as a fast MoE, or a 70B for quality if you are willing to accept about 10 t/s at best (400 ÷ 40).
One point to watch with reasoning models: they produce long blocks of reasoning before the answer, which multiplies the number of tokens to generate. On a machine running at 60 t/s, ten thousand reasoning tokens take nearly three minutes. For simple questions, disable reasoning mode when the model allows it.
#On battery, during extended use: what changes
The MacBook Pro is designed to last longer than an Air, but it is still a laptop. Two settings affect throughput: the power mode (on battery, macOS may limit power) and, on models that offer it, High Power mode. According to Apple, this mode lets the fans run faster to maintain better performance under very intensive workloads, with additional noise. For generation lasting several minutes, connect the power adapter and try this mode.
#Should you upgrade to an M4 Max?
The M4 Max reaches 546 GB/s and has 128 GB of memory according to Apple. In the llama.cpp table, a Llama 7B in Q4_0 goes from 65.95 t/s on a 38-core M2 Max to 83.06 t/s on a 40-core M4 Max: about 26% more. The gain is real for intensive use, but the 96 GB M2 Max retains a capacity advantage over an M4 Pro, which is limited to 64 GB.
| Situation | Recommendation |
|---|---|
| You mainly use 8-14B models | An M2 Pro 32 GB is enough; do not pay extra for bandwidth |
| You run 27–35B models every day | M2 Max 64 or 96 GB: this is where it makes sense |
| You want a comfortable 70B or two resident models | M2 Max 96 GB, or an M4 Max 128 GB if speed matters |
| You’re looking for a headless server | A Mac mini M4 Pro or Mac Studio is better suited |
- MacBook Pro M3
- MacBook Pro M4
- Mac Studio for 64 to 512 GB
- Hardware profile: MacBook Pro M4 Max
- Source: llama.cpp measurements on Apple chips
- Source: Apple, launch of the MacBook Pro M2 Pro and M2 Max
- Source: official Ollama FAQ
#Frequently asked questions
Is a 16 GB MacBook Pro M2 Pro sufficient for local AI?+
What is the difference between the 30-core and 38-core GPU versions of the M2 Max?+
Can you run two LLMs in parallel on a 64 GB M2 Max?+
Which model for French on a MacBook Pro M2 Max?+
Does the M2 Max run hotter than the M1 Max during inference?+
Should you install MLX alongside Ollama on an M2 Max?+
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.