Intermediate 11 minMacBook Pro

Which LLM on MacBook Pro M2 Pro / Max (16–96 GB) ?

Direct response

A MacBook Pro M2 Pro comfortably runs an 8–9B model with 16 GB and a 24–32B model with 32 GB; an M2 Max with 64 or 96 GB is the only variant of this generation that can run a 70B model in Q4 (about 40 GB) and two 20–30B models at the same time. Speed depends on bandwidth: 200 GB/s for the M2 Pro and 400 GB/s for the M2 Max, yielding nearly twice as many tokens per second for the same model.

Released in January 2023, the MacBook Pro M2 Pro / Max remains one of the most capable laptops for local AI: fans, 400 GB/s of bandwidth on the M2 Max, and up to 96 GB of unified memory. This page explains what each memory configuration can handle, what the published measurements show, why doubling bandwidth does not double speed, and when it is worth moving to an M4 Max.

Choosing a machine? Our picks by budget →

By Mohamed Meguedmi·Update 2026-09-30·Tested on macOS 14+
Recommended hardware

Buying alternative for this guide: GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395).

A mini PC is a complete machine: check the available memory and engine compatibility. It does not replace macOS/MLX or CUDA.

Why this choice? Our complete guide on GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395) →

Compare all options by budget, from €800 to €3,500 →

On the go: which laptop for local AI →

Affiliate links — possible commission at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.

#MacBook Pro M2 Pro / Max in 2026: the useful technical specifications

When it was unveiled, Apple announced 200 GB/s of bandwidth and up to 32 GB of unified memory for the M2 Pro, and 400 GB/s, up to 96 GB, and a GPU with up to 38 cores for the M2 Max. Apple presented those 96 GB as pushing the limits of graphics memory in a laptop. The memory is soldered to the processor, so you choose the configuration when you buy it and cannot change it later.

The 14- and 16-inch MacBook Pro chips (2023)
ChipBandwidthPossible memoryGPU
M2 Pro200 GB/s16 or 32 GB16 or 19 cores
M2 Max400 GB/s32, 64, or 96 GB30 or 38 cores

Active cooling is the second advantage over a MacBook Air: the MacBook Pro can sustain long workloads without the chip having to slow down significantly. Apple also offers, on some models, a High Power mode that lets the fans spin faster; it’s worth trying for long generations, at the cost of additional noise.

#M2 Pro or M2 Max: what bandwidth really changes

The Mac Kit

Local AI on your Mac, done right: the models that fit your unified memory, MLX vs Ollama vs LM Studio, and the settings that make Apple Silicon fast.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

Text generation reads all the model’s active weights for every token. Maximum speed is therefore bandwidth divided by weight size. An M2 Max at 400 GB/s could theoretically be twice as fast as an M2 Pro at 200 GB/s. Measurements published in the llama.cpp repository show that this is only partly true.

Text generation, Llama 7B, llama.cpp (discussion #4167)
Chip (GPU cores)BandwidthQ4_0Q8_0F16
M2 Pro (19)200 GB/s38,86 t/s23,01 t/s13,06 t/s
M2 Max (30)400 GB/s60,99 t/s39,97 t/s24,16 t/s
M2 Max (38)400 GB/s65,95 t/s41,83 t/s24,65 t/s

On a 7B model in Q4_0, the M2 Max is only 1.6 to 1.7 times faster than the M2 Pro. On the same model in F16, around 13 GB, the gap rises to nearly 1.9. The calculation explains why: a small model is read so quickly that other limits (compute, latency, overhead) take over, while a large model genuinely saturates memory. The practical consequence is counterintuitive: the larger the model, the more the M2 Max justifies its price. For a 7B, an M2 Pro is more than enough.

→
Order of magnitude for your model
Divide the bandwidth by the Q4 weight size listed in the catalog, then subtract 15 to 35% depending on the model size: the M2 Max reaches about 63% of the ceiling on a 7B in Q4_0 and about 83% on a 7B in F16. For a 19 GB model such as Qwen3-Coder 30B-A3B, the M2 Max ceiling is already 21 t/s for dense reads; MoE models read only their active parameters, so their actual throughput is much better.

#From 16 to 96 GB: what fits and what remains borderline

On Apple Silicon, the GPU does not have access to all memory by default: according to discussions in the llama.cpp repository, Metal grants it between two-thirds and three-quarters. On 16 GB, allow about 10 to 12 GB for the model and its context; on 96 GB, about 64 to 72 GB. The limit can be raised through a system setting detailed in the dedicated guide, but exceeding it may freeze the machine.

Q4 weights (QuelLLM catalog) by memory capacity, excluding context
ModelQ4 weightM2 Pro 16 GBM2 Pro 32 GBM2 Max 64 GBM2 Max 96 GB
Qwen 3.5 9B6 GBComfortableYesYesYes
Gemma 4 12B7 GBJustYesYesYes
Mistral Small 3.2 24B14 GBNoYesYesYes
Qwen 3.5 27B16 GBNoYes, moderate contextYesYes
Qwen3-Coder 30B-A3B (MoE)19 GBNoYes, moderate contextYesYes
Qwen 3.6 35B-A3B (MoE)21 GBNoJustYesYes
Llama 3.3 70B40 GBNoNoYes, short contextYes

MoE models such as Qwen3-Coder 30B-A3B or Qwen 3.6 35B-A3B weigh 19 and 21 GB, but activate only about 3 billion parameters per token: at the same memory capacity, they generate noticeably faster than a 27B dense model. This is the most cost-effective option at 32 GB. At 96 GB, the M2 Max is less useful for loading a single very large model than for keeping two models resident at once, such as a chat model and a coding model.

#Published throughput: what we can and cannot claim

We don't have in-house measurements for this machine, and tokens per second depend on the model, quantization, context, and the version of llama.cpp or Ollama. The only measurements we cite are those published by users on a Llama 7B, which provide the scale. For a newer model, the bandwidth-divided-by-weights calculation remains the right rule of thumb, with the margin of error indicated above.

Prompt processing follows a different pattern: the llama.cpp table reports, for the same Llama 7B in F16, 384 t/s for prompt processing on a 19-core M2 Pro and 756 t/s on a 38-core M2 Max. The number of GPU cores affects this phase, not generation. For RAG or long-document analysis, an M2 Max will therefore wait about half as long for the first word as an M2 Pro.

#Install Ollama and use it correctly on a Mac

Ollama requires macOS Sonoma (14) or later. On Mac, the official FAQ specifies that when Ollama runs as an application, its environment variables must be set with launchctl, and then the application must be restarted. Exports in a terminal are not enough; the application does not see them.

Useful settings for large models
launchctl setenv OLLAMA_FLASH_ATTENTION 1
launchctl setenv OLLAMA_KV_CACHE_TYPE q8_0
# puis quitter et relancer l'application Ollama

The Ollama FAQ states that the q8_0-quantized K/V cache uses roughly half the memory of the default f16 format, and that Flash Attention must be enabled to benefit from it. On a 70B model with a long context, this setting is what makes the cache fit. One limitation to know: cache quantization can slightly reduce quality on some models; test it on your use cases before keeping it.

!
Raise the GPU memory limit only if you understand the implications
The iogpu.wired_limit_mb parameter lets you allocate more memory to the GPU, but it does not survive a reboot, and excessive values can freeze macOS. It is described in the guide to Apple Silicon settings, which remains the reference; apply it only to a model that otherwise will not fit, and leave at least 8 GB for the system.

#Which model for which use case

General-purpose chat
An 8B to 12B model (Qwen 3.5 9B, Gemma 4 12B) responds quickly on an M2 Pro; a 27B model becomes worthwhile starting at 32 GB for more polished French.
Code
A coding MoE such as Qwen3-Coder 30B-A3B on 32 GB or more, or a 24B coding-agent model such as Devstral Small 2 (14 GB in Q4). For online autocompletion, a 7B model is sufficient.
Document RAG
Gemma 4 12B for speed, a 24B model for reasoning quality; limit the context to avoid extending the wait before the first token.
Long analysis, agents
On 64 GB or more: Qwen 3.6 35B-A3B as a fast MoE, or a 70B for quality if you are willing to accept about 10 t/s at best (400 ÷ 40).

One point to watch with reasoning models: they produce long blocks of reasoning before the answer, which multiplies the number of tokens to generate. On a machine running at 60 t/s, ten thousand reasoning tokens take nearly three minutes. For simple questions, disable reasoning mode when the model allows it.

#On battery, during extended use: what changes

The MacBook Pro is designed to last longer than an Air, but it is still a laptop. Two settings affect throughput: the power mode (on battery, macOS may limit power) and, on models that offer it, High Power mode. According to Apple, this mode lets the fans run faster to maintain better performance under very intensive workloads, with additional noise. For generation lasting several minutes, connect the power adapter and try this mode.

#Should you upgrade to an M4 Max?

The M4 Max reaches 546 GB/s and has 128 GB of memory according to Apple. In the llama.cpp table, a Llama 7B in Q4_0 goes from 65.95 t/s on a 38-core M2 Max to 83.06 t/s on a 40-core M4 Max: about 26% more. The gain is real for intensive use, but the 96 GB M2 Max retains a capacity advantage over an M4 Pro, which is limited to 64 GB.

When to keep your M2 Max and when to switch
SituationRecommendation
You mainly use 8-14B modelsAn M2 Pro 32 GB is enough; do not pay extra for bandwidth
You run 27–35B models every dayM2 Max 64 or 96 GB: this is where it makes sense
You want a comfortable 70B or two resident modelsM2 Max 96 GB, or an M4 Max 128 GB if speed matters
You’re looking for a headless serverA Mac mini M4 Pro or Mac Studio is better suited

#Frequently asked questions

FAQ
Is a 16 GB MacBook Pro M2 Pro sufficient for local AI?+
Yes for 8- to 9-billion-parameter models in Q4, which weigh 5 to 6 GB, with a reasonable context. It is not suitable for 24B models and larger. Since memory cannot be added later, choose 32 GB if you want the option to move to a larger model someday.
What is the difference between the 30-core and 38-core GPU versions of the M2 Max?+
Both have 400 GB/s of bandwidth. In llama.cpp benchmarks, text generation is only about 8% faster with 38 cores (65.95 vs. 60.99 t/s in Q4_0) because it is bandwidth-bound. The gap is more apparent in prompt processing, where the 38-core version is significantly faster.
Can you run two LLMs in parallel on a 64 GB M2 Max?+
Yes, if the combined weights and contexts fit within the portion of memory that Metal allocates to the GPU, roughly 43 to 48 GB. For example, a 27B model (16 GB) and a 12B model (7 GB) fit. Ollama keeps models in memory for 5 minutes by default, and reloading takes a few seconds.
Which model for French on a MacBook Pro M2 Max?+
Mistral Small 3.2 24B, from the French Mistral AI lab, is a good candidate for 32 GB and up (14 GB in Q4). Qwen 3.5 9B and Gemma 4 12B are multilingual and faster. Test them on your own texts: French quality depends on the type of task.
Does the M2 Max run hotter than the M1 Max during inference?+
We don't have a reliable numerical comparison to cite. Both have active cooling; speed depends mainly on bandwidth, identical at 400 GB/s. For long generations, connect the power supply and try the High Power mode offered by Apple, accepting the fan noise.
Should you install MLX alongside Ollama on an M2 Max?+
Not necessarily. Ollama is enough for chat, RAG, and the API. MLX appeals to developers who want to control the model from Python or test local LoRA fine-tuning. Speed differences vary by model and version: the site's detailed comparison examines them without promising a fixed gain.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.