Intermediate 14 minApple

MacBook Pro M5 Max 128 GB: real-world local AI measurements (benchmark 2026)

A MacBook Pro M5 Max with 40 GPU cores and 128 GB of unified memory, a protocol written in advance, six models, and a 16,689-token document: here are the first public measurements of the largest Apple laptop for local AI. The measurements were conducted by Joël Ramat, a cybersecurity and GRC consultant, on his own machine, at our request and according to our protocol. All figures are raw outputs from Ollama, with no edits.

Choosing a machine? Our picks by budget →

By Mohamed Meguedmi·Update 2026-09-16·Tested on macOS 14+
Recommended hardware

Buying alternative for this guide: GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395).

Why this choice? Our complete guide on GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395) →

Compare all options by budget, from €800 to €3,500 →

On the go: which laptop for local AI →

Affiliate links — commission possible at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.

MacBook Pro open on a dark desk, terminal displaying text generation in progress
Illustration. The machine tested is a stock 128 GB MacBook Pro M5 Max, with no system settings modified.
i
Why this guide exists
The 128 GB M5 Max is the only high-end configuration in our configurator for which we could not produce any first-hand measurements. Joël Ramat, a reader of the Mac kit and the machine's owner, offered to take the measurements. We wrote the protocol, he ran it on September 16, 2026, and we are publishing his results here, crediting him as the author of the measurements.

#The machine and the protocol

Machine specs: M5 Max chip, 40 GPU cores, 18 CPU cores, 614 GB/s, 128 GB unified memory, macOS 27.0, Ollama 0.34.1
The exact configuration, captured by script before the first measurement. The number of GPU cores is the decisive factor: only the 40-core variant reaches 614 GB/s.

The machine is a 16-inch MacBook Pro with the full M5 Max chip: 18 CPU cores, 40 GPU cores, and 128 GB of unified memory advertised at 614 GB/s. It runs macOS 27.0 with Ollama 0.34.1. No system settings were changed: the GPU memory limit is the default, which matters when comparing it with your own Mac.

The protocol comes down to four rules, and they're not decorative. Keep the Mac plugged in, because macOS throttles the chip on battery power. Close demanding applications. Load only one model at a time. And above all, use the same prompt for every series—the one that makes the numbers comparable across models and, tomorrow, across machines.

The prompt shared by all series
Explique en 300 mots la différence entre la mémoire unifiée d'un Mac et la VRAM d'un GPU dédié, pour un lecteur non technique.
Series A, the two reference points
Gemma 4 12B and Qwen 3.8 27B in GGUF, two models that also run on much more modest Macs. They place the machine on a familiar scale.
B series, MLX versus GGUF
The same Qwen 3.8 27B, with the same weights, served by the MLX engine from Ollama instead of llama.cpp. Only the engine changes.
C series: the real issue at the 128 GB tier
gpt-oss 120B, 65 GB, a model that simply doesn't fit in a laptop PC.
D series, prefill under real-world load
A 16,689-token document sent as a single block to gpt-oss 120B, with a summarization instruction. This is the campaign's only significant prefill measurement.
Two off-protocol bonuses
Joël already had Qwen 3.6 35B-A3B in MLX and gpt-oss 20B on his machine. He measured them under the same conditions.
→
What we measure, and what not to confuse
Ollama displays two throughput figures. Prefill (prompt eval rate) is how quickly the model reads your prompt before responding: it determines the wait before the first word. Generation (eval rate) is the response being produced, token by token: it determines how fluid the experience feels. A model can be fast at one and slow at the other.

#All figures in one table

The Mac Kit

Local AI on your Mac, done right: the models that fit your unified memory, MLX vs Ollama vs LM Studio, and the settings that make Apple Silicon fast.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

Here are the raw outputs of the ollama run --verbose command, as Joël copied them after each measurement. Each model produced a complete response to the 300-word prompt, in reasoning mode when the model supports it, which explains the 1,400 to 3,200 tokens generated for 300 words of final text.

M5 Max 40 GPU · 128 GB · Ollama 0.34.1 · industry · two passes per model, second retained · 16/09/2026
ModelEngineWeightsGenerationGenerated tokensTotal duration
Gemma 4 12BGGUF Q47.6 GB58.1 tok/s1 39024 s
Qwen 3.8 27BGGUF Q417 GB36.6 tok/s2 51769 s
Qwen 3.8 27BMLX18 GB68.0 tok/s2 10331 s
Qwen 3.6 35B-A3B (bonus)MLX23 GB167.2 tok/s3 20519 s
gpt-oss 20B (bonus)MXFP413 GB113.4 tok/s2 19919 s
gpt-oss 120BMXFP465 GB79.1 tok/s6698,5 s
gpt-oss 120B, restart 1MXFP465 GB77.5 tok/s73212,5 s
gpt-oss 120B, rerun 2MXFP465 GB75.5 tok/s1 70722,7 s
gpt-oss 120B, D seriesMXFP465 GB67.6 tok/s2 58450 s
Prefill: only the D-series reading is meaningful (see the Limitations section)
SeriesPrompt tokensPrefillTime to first token
A, B, C (short prompt)43 à 9831 to 346 tok/s0.1 to 0.9 s: startup latency, not throughput
D, 16,689-token document16 6891,388 tok/s12,0 s
!
Why short-prompt prefills aren’t annotated
Each model was measured twice, and only the second pass is retained, in accordance with the protocol: model loading and Metal shader compilation are therefore excluded from the measurement. But on a prompt of 43 to 98 tokens, prefill takes less than a second: the figure reflects fixed startup latency, not reading throughput. It takes thousands of tokens for throughput to take over, and that is precisely the subject of the D series.

#Generation: six models, one read

Bar chart of generation speed: Qwen 3.6 35B-A3B 167 tokens per second, gpt-oss 20B 113, gpt-oss 120B 79, Qwen 3.8 27B MLX 68, Gemma 4 12B 58, Qwen 3.8 27B GGUF 37
Generation in tokens per second. MoE models are shown in orange; they activate only 3 to 5 billion parameters per token. Dense models are shown in blue.

Read naively, this ranking is absurd: a 120-billion-parameter model at 79 tokens per second, twice as fast as a 27B in GGUF. The explanation is three letters. gpt-oss 120B, gpt-oss 20B, and Qwen 3.6 35B-A3B are mixture-of-experts models, or MoE. For each token, only a fraction of the parameters is used: about 5 billion for gpt-oss 120B, 3.6 billion for gpt-oss 20B, and 3 billion for Qwen. Generation speed depends on what moves through memory at each token, not on what is sitting idle there.

Dense models, Gemma 4 12B and Qwen 3.8 27B, use all their parameters for every token. Their speed therefore follows memory bandwidth divided by model weight: this has been the rule for Macs since the M1, and the M5 Max is no exception. On this machine, a 27B dense model in GGUF runs at 37 tokens per second, and a 12B model at 58.

167 tok/s on Qwen 3.6 35B-A3B with MLX
The campaign’s highest score, on the model that combines both accelerators: MoE architecture and the MLX engine. It has the speed of a small 8B model, with the knowledge of a 35B model.
113 tok/s on gpt-oss 20B
The best balance of speed, quality, and size for everyday assistant use: 13 GB, with a 2,200-token response in 19 seconds.
79 tok/s on gpt-oss 120B
The most capable model on the list, running at twice human reading speed. This is the result that justifies the 128 GB tier; we’ll return to it below.
58 and 37 tok/s on dense models
Gemma 4 12B and Qwen 3.8 27B in GGUF: comfortable, but unsurprising. They are in the same ballpark as on an M4 Max, since bandwidth changed little between the two generations.
i
Why the dense 27B is slower than the 120B MoE
Qwen 3.8 27B must move 17 GB of weights per generated token. gpt-oss 120B moves about 3 GB—the weights of the active experts—even though all 65 GB must fit in memory. At 614 GB/s, the former mechanically tops out at around 36 tokens per second, which is exactly what Joël measures. The latter has headroom.

#MLX versus GGUF, all else equal

Comparison of two bars: Qwen 3.8 27B at 36.6 tokens per second in GGUF and 68 tokens per second in MLX, or 86 percent more
Same model, same weight to within one gigabyte; only the engine changes. The measured difference is 86%.

Everyone claims that MLX, the Apple library optimized for Apple Silicon, is faster than llama.cpp. Very few people measure it under strictly identical conditions, using the same model, on the same day, on the same machine. That's what Series B does: Qwen 3.8 27B in GGUF Q4 at 36.6 tokens per second, then Qwen 3.8 27B in MLX at 68.0 tokens per second. 86% faster.

Our guide to MLX and the M5 Mac announced a 30 to 40% gain in generation speed. On the M5 Max and this model, the reality is more than double. Some of the gap likely comes from MLX optimizations specific to the M5 generation and the neural accelerators built into the GPU; we can't yet isolate their contribution. What we can say is that on a recent Mac, running a dense model in GGUF when an MLX version exists is like leaving half the machine in the garage.

Switch a model to the MLX engine
ollama pull qwen3.8:27b-mlx
ollama run qwen3.8:27b-mlx --verbose
→
The habit to develop
Before pulling a model, check whether a -mlx tag exists in the Ollama library. For a dense model of 20 GB or more, the difference is immediately visible. The guide to MLX and llama.cpp explains how the engine is selected and what to do when an MLX tag is missing.

#What 128 GB actually buys you

Diagram: on the left, unified memory, a single 128 GB pool shared between CPU and GPU that accommodates the 65 GB AI model; on the right, dedicated memory, where 65 GB does not fit in 16 GB of VRAM
Conceptual diagram. Unified memory makes the entire 128 GB pool available to the GPU. On a laptop PC, a 16 GB graphics card cannot load a 65 GB model, regardless of system RAM.

The real issue with the 128 GB tier is not speed, but what fits. gpt-oss 120B weighs 65 GB in MXFP4. No PC laptop can run it on the GPU: mobile cards top out at 16 or 24 GB of VRAM. On the 128 GB M5 Max, it loads, responds at between 75 and 79 tokens per second across three runs, and still leaves room.

Stacked bar for the 128 GB: 35 GB for macOS and open applications, 65 GB for gpt-oss 120B, 28 GB free
Memory distribution during the C series, in orders of magnitude. Joël closed nothing: Obsidian, Terminal, and the rest of his session were open.

The memory report taken before the run is telling: 35 GB occupied by macOS and the work-session applications, with 93 GB available for AI. The 120B fits with about thirty gigabytes to spare, enough to handle a long context and a second lightweight model in parallel, such as an 8B for code autocompletion. On an M5 Max 64 GB, the same model does not load at all; on 96 GB it runs, but with no room left for anything else.

8.5 to 22.7 seconds
Total duration of gpt-oss 120B's response to the 300-word prompt, over three runs: 669, 732, then 1 707 tokens produced. The speed barely changes (79.1, 77.5, and 75.5 tokens per second); the reasoning length varies from one run to the next, depending on whether the model decides to count its words one by one.
50 seconds
Series D duration: reading a 16,689-token document, then writing a ten-bullet summary of 2,584 tokens.
“Nominal” thermal pressure
Measured after each series. No audible fan noise throughout the entire campaign, including 120B.
i
What we couldn't measure at this tier
A dense 70B in Q4—the configuration our MacBook Pro M5 Max spec sheet estimates at between 15 and 25 tokens per second in MLX—was not installed on the machine. The figure remains an estimate until the next test run.

#Prefill, the number nobody publishes

Timeline for a 50-second request: 12 seconds to read a 16,689-token document at 1,388 tokens per second, followed by 38 seconds of response generation
Series D. The document sent is the benchmark report itself, exported as text, followed by a summarization instruction.

A twenty-word prompt says nothing about prefill. To measure it, Joël sent gpt-oss 120B a text of 11 280 words in a single block—16 689 tokens and about 45 pages—with a summary instruction. Result: 1 388 tokens per second while reading, meaning twelve seconds before the first response word appeared, followed by a 2 584-token summary at 67.6 tokens per second.

This is the number that matters for serious use cases. Analyzing a contract, summarizing an audit report, querying a document database with RAG: in all these cases, the model spends most of its time reading, not writing. Twelve seconds for 45 pages on a laptop, without a single byte leaving the machine—that’s what makes local AI usable for a consultant bound by professional confidentiality or a company subject to strict confidentiality requirements.

→
Rough sizing for your documents
At 1,388 tokens per second, a 30-page contract is read in 8 seconds, and a 100-page report in 27 seconds. These times are added to response generation, so allow about one minute for a complete summary of a long document on this machine.

#Heat and fans

Joël recorded the thermal state after each series using macOS’s thermal stress command. The verdict was consistently “Nominal,” with no audible fan noise, including during series D, which drives the GPU at full load for fifty seconds. This is consistent with what we observe on the M4 Max, with one caveat: our measurements are bursts lasting less than a minute. The throttling test specified in the protocol—ten minutes of continuous load followed by a new measurement of series A—was not performed. We will request it during the next campaign.

Check thermal pressure during generation
sudo powermetrics --samplers thermal -i 1000 -n 1 | grep -i pressure

#What these measurements don’t tell you

An honest benchmark lists what it does not prove. Here are the caveats, in order of importance.

Short prompt prefills
Two passes per model, with the second retained. But 43 to 98 prompt tokens measure only startup latency: only series D, at 16 689 tokens, provides a true reading throughput.
No 70-billion-parameter dense model
The 128 GB tier is also justified by 70B models at Q4 with a long context. Not measured here.
No throttling test
Bursts of less than a minute. Behavior over an hour of intensive RAG remains to be documented.
No battery measurement
The protocol requires mains power. On battery, expect significantly lower figures, as macOS throttles the chip.
One clean Series D document for each tester
The long text sent in Series D is not mandated by the protocol: two testers do not read the same document, which limits comparisons of prefills between machines. At Joël's suggestion, the protocol now provides a single 11,292-word reference text to download in the next section: future measurements will be comparable. The measurement on this page was taken earlier, using another document of equivalent length.
A single machine, a single operator
No variance measured. If you have the same configuration, your measurements are welcome; see the next section.
Reasoning-mode models
Models that think before answering produce far more tokens than the requested 300 words, and this length varies from one run to another for the same model (669 to 1 707 tokens on gpt-oss 120B). Total durations are therefore not comparable; only speeds are.
Prompt caching on retries
When the same prompt is run again, Ollama reuses the tokens it has already read (“prompt eval cached”). The prefill of an identical rerun is therefore not a measurement; to measure prefill, you need a long, new prompt, as in the D series.

#Reproduce the benchmark at home

Laptop viewed from above on a white desk, mechanical stopwatch, charger plugged in, notebook and pen
Illustration. Connected sector, closed applications, one model at a time: three conditions without which a benchmark is meaningless.

The protocol is designed to be replayed as-is on any Mac Apple Silicon, from MacBook Air to Mac Studio. If you run the same series with the same prompt, your figures are directly comparable to those on this page.

  1. 01
    Prepare the machine
    Plug in the power cable, close your browser, Docker, and virtual machines. Note your chip model, GPU core count, and Ollama version.
  2. 02
    Check the system status
    Record the chip, number of GPU cores, macOS version and Ollama, power supply, available memory, and open applications. Joël wrote a zsh script that generates this report with one command; it will be added to the Mac kit’s benchmark chapter after verification, with his permission.
  3. 03
    A series
    Pull gemma4:12b and qwen3.8:27b, run each with --verbose, paste the shared prompt, let the response finish, and exit with /bye. Run two passes and keep the second.
  4. 04
    Series B
    Same with qwen3.8:27b-mlx. Compare the eval rate line with the one for the A series.
  5. 05
    Series C
    If you have 96 GB or more: gpt-oss:120b. Otherwise, choose the largest model that fits in your unified memory minus 10 GB.
  6. 06
    D Series
    Download the QuelLLM reference text, 11,292 words, the same for everyone, and pass it to ollama run followed by the instruction “Summarize this text in 10 bullet points.” Note the displayed prompt-token count: it depends on the model, not the text.
Protocol commands
# Série A — les deux repères
ollama pull gemma4:12b && ollama pull qwen3.8:27b
ollama run gemma4:12b --verbose
ollama run qwen3.8:27b --verbose

# Série B — même modèle, moteur MLX
ollama pull qwen3.8:27b-mlx
ollama run qwen3.8:27b-mlx --verbose

# Série C — le palier 128 Go (65 Go à télécharger)
ollama pull gpt-oss:120b
ollama run gpt-oss:120b --verbose

# Série D — préfill sur le texte de référence commun (11 292 mots)
curl -sO https://quelllm.fr/img/protocole/texte-reference-quelllm-v1.txt
ollama run gpt-oss:120b --verbose "$(cat texte-reference-quelllm-v1.txt) Résume ce texte en 10 puces."
→
Send us your reports
Paste the raw statistics block from each measurement into an email to contact@quelllm.fr, along with your chip, GPU cores, and Ollama version. We publish verified results with your name, or anonymously if you prefer. The configurations we are missing most: M5 Max 64 GB, M5 Pro, Mac Studio M5 Ultra.

#Who it's for, and at what price

The 128 GB M5 Max is justified for one use case: running locally models that fit nowhere else in a laptop, with room for context and a second model. gpt-oss 120B at 79 tokens per second and a 45-page document read in twelve seconds, without a fan, is an AI workstation that fits in a bag.

Reading results by usage profile
You areWhat this benchmark tells youRecommended tier
Consultant, lawyer, or certified public accountant under professional secrecyA 120B model reads and summarizes your files locally, at a usable speed, without sending anything outsideM5 Max 128 GB
Developer who wants a serious local copilotgpt-oss 20B or Qwen 3.6 35B-A3B are more than sufficient and fit within 48 GB. Our 113 and 167 tok/s figures apply to the 40-core GPU chip: there is enough memory, but speed will be lower on a variant with lower bandwidthM5 Max 48 or 64 GB
Daily user of a writing assistantA dense 12B or a 20B MoE covers the use case; 128 GB is a luxury. There is enough memory, but the M5 Pro's is slower: expect less than our 58 and 113 tok/sM5 Pro 48 GB
Team that wants a shared serverThe Mac Studio M5 Max or Ultra offers the same engine with more memory and desktop coolingMac Studio
!
These speeds apply to this chip
All figures on this page were measured on an M5 Max with 40 GPU cores and 614 GB/s of memory bandwidth. LLM generation is limited by this bandwidth: on an M5 Max with 32 GPU cores or an M5 Pro, whose memory is slower, the same models will fit in memory but run more slowly. The table above tells you what fits, not the speed you will get on another tier.

If your needs stop at 30-billion-parameter models, half the memory is enough, and the price difference can fund a very good monitor. If your needs start at 100-billion-parameter models, there is no portable alternative, and this guide gives you the figures to justify that decision to whoever signs the purchase order.

#Credits and sources

Consultant working at night on a laptop displaying a terminal, with a desk lamp and notebook
Illustration. An ordinary work session with applications open: these are the conditions under which the measurements were taken, on the author's work machine.

The measurements on this page were taken by Joël Ramat, a cybersecurity and GRC consultant, ISO 27001 expert, and advisor and auditor specializing in AI-augmented on-premises services, as well as cofounder and chairman of Sywédgia SAS. He ran the QuelLLM protocol on his own machine on September 16, 2026, and reviewed this article before publication. The measurements remain his, and he is free to publish them separately.

Method: raw outputs from ollama run --verbose copied after each measurement, with no editing; system state recorded by script before the test run; thermal throttling read after each series. The full report, including the models’ complete responses, is retained and can be provided upon request. Written and formatted by Mohamed Meguedmi, BestLLMfor.


#FAQ

Is the 128 GB M5 Max faster than the 128 GB M4 Max for local AI?+
Not dramatically on dense GGUF models: memory bandwidth has barely changed, and Qwen 3.8 27B runs at 37 tokens per second, roughly comparable to the M4 Max. The gap widens with MLX, where the M5 Max measures 68 tokens per second on the same model. We have no M4 Max MLX measurement under the same conditions to quantify the exact gap.
Why is gpt-oss 120B faster than Qwen 3.8 27B?+
Because gpt-oss 120B is an MoE model that activates only around 5 billion parameters per token, whereas Qwen 3.8 27B, a dense model, moves 27 billion at every token. Speed depends on active parameters, not total parameters. However, the 65 GB required by the 120B must fit in memory, which only the 96 or 128 GB tier allows.
Do you need 128 GB to run gpt-oss 120B?+
You need at least 96 GB of unified memory to load it with a reasonable context. The 128 GB version provides room for a long context, a second lightweight model, and a normal work session open alongside it. Joël ran the 120B with 35 GB already occupied by his applications.
Do these figures hold on battery power?+
No. The protocol requires the sector because macOS throttles the chip on battery power. Expect significantly lower speeds on the go, and especially measurements that cannot be reproduced.
Why aren’t the prefills for series A, B, and C published?+
They were measured on prompts ranging from 43 to 98 tokens. At that size, prefill takes less than a second and reflects startup latency, not reading throughput; the figure says nothing about how fast the model reads a real document. The only usable prefill figure is from series D, on 16,689 tokens: 1,388 tokens per second.
Can I submit my own benchmarks?+
Yes. Follow series A through D with the common prompt, two passes per measurement, and send the raw statistics blocks to contact@quelllm.fr with your configuration. We publish the verified results and credit you.
Where can I find the preparation script used for this benchmark?+
The zsh script written by Joël Ramat, which generates a complete report on the Mac's state before measurement, will be added to the Mac kit's benchmark chapter once it has been verified on other machines, with his permission and attribution. In the meantime, the “Reproducing the benchmark” section lists the information to record manually.

Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.

Prices in euros (€) are French market prices including VAT, as checked by BestLLMfor. US prices differ: the Amazon buttons show the current US price.