Intermediate 11 minApple

Local LLM on Mac M5: MLX now outperforms llama.cpp (figures 2026)

On Mac M5, the MLX-versus-llama.cpp debate is settled: with the same model and quantization, the Apple framework now delivers 30 to 40% more tokens per second. This guide quantifies the gap on M5 and M5 Max, explains the new MLX engine from Ollama released in June 2026, and shows how far you can push things—including linking two machines over Thunderbolt 5 to run 120B+ models. It is the M5 companion to the MacBook Pro M4 Max guide: same principles, updated figures.

Choosing a machine? Our picks by budget →

By Mohamed Meguedmi·Update 2026-08-30·Tested on macOS 14+
Recommended hardware

Buying alternative for this guide: GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395).

Why this choice? Our complete guide on GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395) →

Compare all options by budget, from €800 to €3,500 →

On the go: which laptop for local AI →

Affiliate links — possible commission at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.

#Why the M5 changes the game for a local LLM on Mac

The M5 chip brings two things that really matter for inference: GPU cores with matrix-multiplication units (Neural Accelerators) integrated into each core, and higher memory bandwidth. LLM inference is primarily a memory-bandwidth problem—each generated token rereads the model's entire weight set. The faster the unified memory, the faster tokens are produced.

This is precisely where Apple's MLX framework has an advantage over llama.cpp. MLX is written to use these new GPU matrix units on the M5 directly, whereas llama.cpp's Metal path remains more generic. The concrete result: on the same machine, with the same model and quantization, MLX generates significantly more tokens per second. The gap, marginal on M3, becomes clear on M5.

i
This guide supplements; it does not replace
If you're using a MacBook Pro M4 / M4 Max, your dedicated guide remains the reference. Here we focus on what the M5 changes: the new figures, the MLX engine from Ollama, and Thunderbolt 5 clusters. The fundamentals (unified memory, quantization choices) are identical from one generation to the next.

#M5 and M5 Max benchmarks: the 2026 figures

The Mac Kit

Local AI on your Mac, done right: the models that fit your unified memory, MLX vs Ollama vs LM Studio, and the settings that make Apple Silicon fast.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

Here are approximate figures measured during generation (decode) with MLX, on common models using 4-bit quantization. As always with local inference, these rates vary with context length, the exact model version, and the machine's thermal conditions: treat them as reference points, not guarantees accurate to the token.

→
Real-world measurements on an M5 Max 128 GB
Since September 2026, we have published a comprehensive benchmark run on a MacBook Pro M5 Max with 40 GPU cores and 128 GB: six models, MLX versus GGUF on the same Qwen 3.8 27B, gpt-oss 120B, and a prefill measured on an 11,280-word document. The orders of magnitude below remain the original ones; the measured figures are in the dedicated guide.
M5 (10-core GPU)—8B Q4
~55-70 tokens/s during generation. An 8B model (Qwen 3.5, Granite 4.2, Gemma 4) remains very comfortable on the base chip.
M5 Max — 8B Q4
~230 tokens/s. The Max's higher memory bandwidth makes the difference on small models, where the GPU is never the bottleneck.
M5 Max — 32B Q4
~55–65 tokens/s. A 32B (Qwen 3.8 27B, Devstral) runs at a comfortable reading speed and is highly usable interactively.
M5 Max — 70B Q4
~28 tokens/s. A 70B in Q4 fits in memory with just 48-64 GB of unified RAM and remains fluid for chat and code generation.

The striking point is 70B at ~28 tokens/s on a machine with no dedicated graphics card, one that's quiet and fits in a bag. For comparison, you need a RTX 4090 24 GB (and partial CPU offloading, because a 70B Q4 weighs ~40 GB and doesn't fit in 24 GB of VRAM) to compete on PC — with much more noise and power consumption.

→
Unified RAM is your VRAM
On Mac, there is no separate VRAM: the model loads into unified memory shared with the system. Q4 size guide: 8B≈5 GB, 32B≈19 GB, 70B≈40 GB. Always allow ~20% headroom for context and macOS. A 70B Q4 therefore requires at least 48 GB, or 64 GB for comfortable operation.

#MLX vs. llama.cpp: where the gap widens

The 30–40% gap in favor of MLX isn’t uniform: it depends on the workload profile. Understanding where it widens helps you know when switching to MLX is really worthwhile.

Generation (decode)
This is where MLX has its clearest advantage on M5, thanks to direct use of the GPU's matrix units. For interactive chat, this is the difference you feel.
Prompt processing (prefill)
MLX retains the advantage, but the gap is somewhat narrower. With very long contexts, both engines become memory-limited.
Model support
llama.cpp remains more universal through the GGUF format; MLX requires a version converted to the MLX format. For popular models, the conversion almost always exists (the mlx-community hub on Hugging Face).
Python integration
MLX is native to Python, making it the obvious choice for prototyping, light fine-tuning, or a custom pipeline. llama.cpp uses bindings.
!
MLX is not magic when it comes to memory
MLX speeds up inference; it does not reduce the model's memory footprint. A model that does not fit in your unified RAM will not fit any better under MLX. The quantization choice (Q4_K_M by default) remains the No. n°1 lever for staying within your memory budget.

#Prerequisites and recommended RAM by model

Before installing, calibrate the model size to your unified memory. Here are the realistic tiers on M5 and M5 Max, in Q4.

16 GB (M5)
3B to 8B models are comfortable in Q4. A 14B model works but leaves little room for a large context.
24–32 GB (M5 / M5 Pro)
A 14B handles comfortably; a 32B Q4 is workable at the upper end. The sweet spot for versatile everyday use.
48 GB (M5 Max)
A 32B Q4 is very comfortable, while a 70B Q4 is barely adequate — it's better to target 64 GB for a 70B with context.
64–128 GB (M5 Max)
A smooth 70B Q4 with a large context, or several models loaded in parallel. Power-user territory.

On the software side, you need up-to-date macOS and one of two paths to MLX: LM Studio (which automatically switches to MLX when an MLX version of the model exists) or Ollama since its June 2026 update, detailed just below.

#Ollama’s MLX engine (June 2026)

Historically, Ollama relied on its own engine and llama.cpp, so it used the Metal path on Mac. Since its June 2026 update, Ollama can run models through MLX on Apple Silicon when an MLX version is available—finally bringing the most widely used local AI tool in line with the performance that LM Studio already delivered through MLX.

In practice, this means you get the MLX performance boost without changing your habits: the same Ollama daemon on the default port 11434, the same commands, and the same OpenAI-compatible API. The engine chooses MLX when appropriate and falls back to the standard path otherwise.

i
Check your Ollama version
MLX support requires an Ollama version from June 2026 or later. Update before comparing throughput; otherwise, you'll still be measuring the old Metal path and wrongly conclude that “MLX changes nothing.”

#Install and run a model in MLX

Two options, depending on your comfort with the terminal. The most direct way to test MLX without configuration is LM Studio; the most integrable into a stack is Ollama.

  1. 01
    1. Update the tool
    Install the latest version of Ollama (June 2026 or newer) or LM Studio. This is required to use the MLX path.
  2. 02
    2. Choose a model with an MLX version
    On LM Studio, search prioritizes MLX variants on Apple Silicon. On Ollama, pull a current model: the engine switches to MLX when a converted version exists.
  3. 03
    3. Launch and verify throughput
    Ask a long question and watch the displayed tokens/s. Optionally compare it with the same model in GGUF to measure the actual difference on your machine.
  4. 04
    4. Adjust quantization if needed
    If the model fills your unified memory, step down a level (Q5 → Q4_K_M) rather than switching tools: memory is the limiting factor, not the engine.
Terminal
# Vérifier la version d'Ollama (le support MLX date de juin 2026)
ollama --version

# Lancer un modèle : Ollama emprunte MLX quand une version MLX existe
ollama run qwen3:8b

# Le daemon expose l'API compatible OpenAI sur le port par défaut
curl http://localhost:11434/api/tags

On the Python side, MLX is directly suited to scripting, which is useful for benchmarking or integrating inference into a custom pipeline:

Terminal
# Installer les outils MLX pour LLM (framework Apple, natif Python)
pip install mlx-lm

# Générer avec un modèle converti au format MLX (hub mlx-community)
mlx_lm.generate --model mlx-community/Qwen3-8B-4bit \
  --prompt "Explique la mémoire unifiée du M5 en trois phrases"

#Thunderbolt 5 clusters for 120B+ models

A single Mac, even a well-equipped M5 Max, is limited by its unified memory. To go beyond that — run a 120B+ model or a large MoE — the 2026 approach is to connect multiple Apple Silicon machines over Thunderbolt 5 and split the model across them. Thunderbolt 5 offers significantly more bandwidth than Thunderbolt 4, finally making activation exchange between nodes viable for inference.

The principle is simple: the model is split into layers, each machine hosts part of the weights in memory, and activations move from one node to the next at each token. This lets you combine the unified memory of several Macs to load a model that would fit on none of them individually.

What this unlocks
120B+ models in Q4, or even very large MoE models, by combining, for example, two 128 GB M5 Max systems to approach 256 GB of addressable memory.
The trade-off
Inter-node latency costs tokens/s: a cluster is slower than a single machine that could fit the same model. You use one because no single machine is sufficient, not to gain speed.
The wiring
Thunderbolt 5 is the key: its bandwidth makes activation transfer acceptable. With Thunderbolt 4 or Ethernet, the interconnect becomes the bottleneck again.
!
The cluster remains a niche use case
Linking Macs for a 120B+ is impressive but reserved for a genuine need: the complexity, the cost of two machines, and the throughput loss are justified only if you cannot simply choose a smaller model or a Mac Studio Ultra with lots of memory. For most use cases, a single M5 Max is more than sufficient.

#Tips and troubleshooting

“MLX doesn't run faster”
First verify that you are actually using the MLX path (up-to-date tool version, with the model's MLX variant really loaded). A GGUF loaded through the Metal path does not benefit from MLX.
Throughput collapses after a few minutes
This is thermal throttling, especially on MacBook Air (which has no fan). Under sustained load, a MacBook Pro or Mac mini/Studio maintains a more stable throughput.
Model that refuses to load
Unified memory is full. Drop down one model size or one quantization level; close RAM-intensive applications.
Long context that crawls
Prefilling a very long prompt is memory-intensive. Reduce the context window if you don't need it, or accept a slower first token.

#Go further

These guides naturally build on what you just read, from the in-depth comparison to hands-on implementation:

MLX vs. llama.cpp in detail
“MLX vs llama.cpp on M-series Macs: who wins in 2026?” digs deeper into the underlying comparison (model support, quantization, Python integration) beyond the M5 figures alone.
Install Ollama on macOS
“Install Ollama on macOS (Apple Silicon)” covers step-by-step installation and use of unified memory, an essential foundation before targeting MLX.
Choose your quantization
“Choosing your quantization (Q4, Q5, Q8, FP16)” explains the quality/memory tradeoff — the key lever for fitting a model into your unified RAM.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.