Intermediate 15 minMac

M4 Max Mac and local LLM: MLX, performance, models

The MacBook Pro M4 Max (64 or 128 GB of unified memory) has become the most capable portable machine for running a local LLM. The combination of unified memory and Apple's MLX framework makes it possible to load today's largest open-weight models—the 30–35B MoEs at full precision and dense 27B multimodal models—on a laptop, at surprisingly fast speeds. This guide measures what actually works: MLX vs. Ollama, practically usable models, measured tokens/sec, and thermals on battery.

Choosing a machine? Our picks by budget →

By Mohamed Meguedmi·Update 2026-08-27·Tested on Windows, macOS, and Linux
Recommended hardware

Buying alternative for this guide: BOSGAME M5 128GB / 2TB (Ryzen AI Max+ 395).

Why this choice? Our complete guide on BOSGAME M5 128GB / 2TB (Ryzen AI Max+ 395) →

Compare all options by budget, from €800 to €3,500 →

On the go: which laptop for local AI →

Affiliate links — possible commission at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.

#Why the M4 Max changes the game for local LLMs

On a Windows or Linux PC, the GPU's VRAM is the main constraint. A RTX 4090 is capped at 24 GB: running a 30-35B model locally in FP16 requires CPU offloading, which cuts throughput by a factor of five or ten. The M4 Max approaches the problem from the other direction. The GPU and CPU share the same memory: everything in RAM is accessible to the GPU without copying.

In practical terms, a 128 GB M4 Max MacBook Pro can load the best 2026 MoE models — Qwen 3.6 35B-A3B or Qwen 3.8 27B — at full FP16 precision, or several models at once, without breaking a sweat, and keep running them on battery. No portable x86 machine can do that today — you need a desktop with a RTX 5090 or two RTX 3090 to approach the same result, and even then, not on the go.

i
What “M4 Max” means
The M4 Max is the high-end chip in the 14" and 16" MacBook Pro models released in late 2024. It comes in two GPU configurations (32 or 40 cores) and four RAM tiers: 36, 48, 64, or 128 GB. For LLM inference, only the 64 and 128 GB versions make sense. The advertised memory bandwidth is 546 GB/s.

#Unified memory: why 128 GB ≠ 128 GB of VRAM

The Mac Kit

You've seen what an M4 Max with MLX can do. The Mac Kit explains when MLX really beats GGUF (ch. 3), how to calculate your memory budget for long contexts (ch. 5), and what performance to expect, including heat and battery life (ch. 6).

  • Lifetime online access
  • PDF + files
  • Lifetime updates

macOS does not let a single process consume all the memory. By default, the GPU can use about 75% of the total RAM. On a 128 GB machine, that gives you ~96 GB available for your models. On a 64 GB machine, you get ~48 GB. That is more than enough for the target models, but you need to know this before calculating.

Mac M4 Max 64 GB
About 48 GB for the LLM. The best 2026 MoE models in Q8 fit comfortably—Qwen3-Coder 30B-A3B (≈32 GB), Qwen 3.8 27B (≈30 GB), Qwen 3.6 35B-A3B (≈40 GB), which just fits—and Mistral Small 24B even fits in FP16 (≈48 GB).
Mac M4 Max 128 GB
About 96 GB for the LLM. Room to run large MoE models at full FP16 precision (Qwen 3.6 35B-A3B, Qwen 3.8 27B, Granite 4.2 30B), keep multiple models loaded in parallel, or push the context to 256k tokens.
Recommended quantizations
Q4_K_M remains the right default. With 128 GB, you can move up to Q5_K_M or Q6 without difficulty for better quality.
Increase the memory allocated to the GPU (optional)
# Par défaut macOS laisse ~75% de la RAM au GPU.
# Pour pousser à ~85% (à vos risques, peut faire swapper le système) :
sudo sysctl iogpu.wired_limit_mb=110000

# Pour revenir à la valeur par défaut :
sudo sysctl iogpu.wired_limit_mb=0
!
Do not push beyond 90%
Setting the limit too high starves macOS itself. Above 90%, you risk system freezes and constant swapping. The default value is suitable for 95% of use cases.

#MLX vs Ollama: which one for an M4 Max Mac?

Apple released MLX in December 2023: an array-computing framework designed for Apple Silicon, functionally equivalent to PyTorch but built around unified memory. The mlx-lm module provides a CLI and Python API for running models quantized in the MLX format (similar to GGUF but distinct).

Ollama (Metal)
Easier to install, a huge ecosystem, universal GGUF formats, and an OpenAI-compatible API on :11434. Slight overhead, with suboptimal CPU↔GPU memory conversion.
MLX (mlx-lm)
Faster in practice on Apple Silicon: +20 to +40% tokens/sec on 30B+ models, native memory management, built-in prompt caching. Smaller ecosystem, with no standardized REST API out of the box.
Verdict
To explore and try 10 models: Ollama. To fully leverage the M4 Max and run large 30–35B MoEs at full precision in local production: MLX.
→
The two coexist very well
Many M4 Max users keep Ollama as a permanent server for third-party apps (Open WebUI, n8n, Continue.dev) and launch MLX occasionally for sessions where they want maximum tokens/sec.

#Install MLX and run a first model

  1. 01
    Prepare a Python environment
    Use Python 3.10+ (3.12 recommended). On Mac, uv or pyenv is the cleanest option. mlx-lm has no native prerequisites other than Apple Silicon.
  2. 02
    Install mlx-lm
    One pip command is all it takes. Everything is compiled for arm64 Metal.
  3. 03
    Run an MLX model
    The mlx_lm.generate CLI downloads the model from Hugging Face (mlx-community organization) and runs it.
  4. 04
    Expose an HTTP API (optional)
    Since version 0.18, mlx-lm has included an mlx_lm.server command that exposes an OpenAI-compatible API on port 8080.
Terminal — installation and first test
# 1. Environnement
python3 -m venv ~/mlx-env
source ~/mlx-env/bin/activate

# 2. Installation
pip install --upgrade mlx-lm

# 3. Téléchargement + génération en une commande
mlx_lm.generate \
  --model mlx-community/Qwen3.6-35B-A3B-Instruct-4bit \
  --prompt "Explique en français le principe de la mémoire unifiée Apple Silicon." \
  --max-tokens 512
OpenAI-compatible MLX server
# Démarre une API sur http://localhost:8080/v1
mlx_lm.server \
  --model mlx-community/Qwen3.8-27B-Instruct-4bit \
  --port 8080

#Models tested on a MacBook Pro M4 Max 128 GB

All figures below were measured on a 16-inch M4 Max MacBook Pro (40 GPU cores, 128 GB), plugged into power, first cold and then after 5 minutes of warm-up. Short prompt (50 tokens), 500-token generation, batch 1.

Qwen 3.6 35B-A3B Q8 (MLX)
≈40 GB in RAM. 42–50 tokens/s: the MoE activates only 3B parameters, hence the high throughput despite its size. The reliable all-purpose choice on this Mac.
Qwen 3.8 27B Q8 (MLX)
≈30 GB of RAM. 18-22 tokens/s (dense model). 262k context and vision, the closest thing to a local Copilot. Consider setting its reasoning level to low; otherwise, it overthinks.
Qwen3-Coder 30B-A3B Q8 (MLX)
≈32 GB of RAM. 44–52 tokens/s. Code-specialized MoE, 256k context: ideal as a local development assistant.
GLM 4.7 Flash Q8 (MLX)
≈32 GB of RAM (MoE 30B-A3B, MIT license). 40–48 tokens/s. Excellent for agents and tool calls, highly responsive.
Granite 4.2 30B Q8 (MLX)
≈33 GB of RAM. 30–38 tokens/s. Geared toward professional/enterprise use, economical with tokens and stable during long sessions.
Mistral Small 24B FP16 (MLX)
≈48 GB. 18–22 tokens/s. Dense model with excellent French quality, particularly good at document summarization.

#Detailed benchmarks: MLX vs Ollama vs llama.cpp

For the same Qwen 3.8 27B model in Q8, here is the spread observed among the three runtimes on the same M4 Max 128 GB machine:

8-bit MLX (mlx-community)
20.5 tokens/s on average, prompt eval ~410 tokens/s.
Ollama Q8_0 (Metal)
15.2 tokens/s on average, prompt eval ~300 tokens/s.
Pure llama.cpp Q8_0
15.6 tokens/s on average. Ollama adds little overhead compared with llama.cpp.
i
Why MLX wins
MLX avoids conversion between CPU and GPU representations and makes better use of Metal Performance Shaders instructions for quantized matmul operations. The difference widens as the model size increases: on 7B models the gap is marginal, while on large 30B+ models it becomes clear.

#M4 Max vs RTX 4090: who wins at what

The candid comparison: a 128 GB MacBook Pro M4 Max (model replaced by the M5 Max at Apple — look for it used, with variable pricing) versus a desktop equipped with a used RTX 4090 24 GB (variable pricing, GPU only, without the rest of the tower).

Small models ≤14B Q4 (Qwen 3.5 9B, gpt-oss 20B)
RTX 4090 wins decisively (80-120 tok/s vs 35-60 on M4 Max). For these sizes, choose the 4090 if you have the choice.
Dense 24–30B Q4 models (Mistral Small, Qwen 3.8 27B)
RTX 4090 still ahead (30-40 tok/s vs. 18-22 on M4 Max), but the gap is narrowing.
Full precision Q8/FP16
The M4 Max pulls ahead in practice: it runs the best 2026 MoE models (35B-A3B, 27B) in Q8 or FP16 within its RAM, whereas the RTX 4090 has to drop to Q4 or offload to the CPU (throughput divided by five).
Large MoE (Qwen 3.6 35B-A3B)
The M4 Max is excellent thanks to its massive RAM: everything fits in memory, even in FP16. RTX 4090 must offload as soon as you exceed Q4.
Mobility
No match: the M4 Max gets ~3 hours of continuous inference on battery, while the RTX 4090 goes 0 km.
→
The decision threshold
Want to run the best 2026 MoE models at full precision (Q8/FP16) and keep several loaded at once? The M4 Max 128 GB is the most cost-effective portable machine on the market. Targeting very fast 13B? A fixed RTX 4090 costs half as much and runs twice as fast.

#Thermals, fans, and battery life

The M4 Max runs relatively cool during inference compared with a NVIDIA GPU. Under sustained load (5 minutes of generation on a large 35B MoE in Q8), the fans ramp up to about 2 800 RPM—audible but far from the noise of a gaming PC. CPU/GPU temperature levels off at around 95 °C without significant throttling on AC power.

Plugged in (140 W charger)
Full power, maximum throughput. No limit, moderate fan speed.
On battery power
macOS slightly reduces the GPU frequency: ~80% of sector throughput. Qwen 3.6 35B-A3B goes from about 45 to 37 tokens/s.
Inference battery life
16-inch MacBook Pro M4 Max, 100 Wh: ~3 hours of continuous inference on a large 35B MoE, ~5 hours on 7B–14B models, ~8 hours of mixed chat + reading.
Fan noise
Audible after 30 seconds of continuous generation, but much quieter than an RTX PC laptop.

#Tips for getting the most out of the M4 Max

Keep the model loaded
The first prompt after loading is slow. Use server mode (mlx_lm.server or ollama serve) to keep the model in RAM for the entire session.
Prefer the MLX format
Models published by mlx-community on Hugging Face are optimized for Apple Silicon. Look for the "mlx-community/" prefix before anything else.
Monitor with asitop or Stats
asitop (the nvtop equivalent for Apple Silicon) shows GPU usage, memory, and instantaneous power draw in real time.
High-performance mode
In System Settings → Battery → select “High Power Mode.” The gain is marginal but real for 70B models.
Disable Spotlight during benchmarks
Spotlight indexing can consume 5 to 10% of throughput. mdutil -a -i off disables it for the duration of a session.
Install asitop to monitor the GPU
# Installation
pip install asitop

# Lancement (nécessite sudo pour lire les compteurs hardware)
sudo asitop
!
No CUDA = no serious fine-tuning
MLX supports basic LoRA fine-tuning, but its ecosystem still lags far behind PyTorch + CUDA. For intensive fine-tuning, an M4 Max Mac isn't the right machine—prefer a RTX 3090/4090 or an H100 cloud service.

#Go further

Three ways to explore the topic further:

Install Ollama on macOS
If you want simplicity before maximum performance, the guide "Install Ollama on macOS (Apple Silicon)" covers all the Metal tooling on the Ollama side.
Choose your quantization
With 128 GB of unified RAM, you have enough headroom to move up to Q5_K_M, Q6, or even FP16—the “Choosing your quantization” guide details the trade-offs.
Mac Studio Ultra for going further
If 30–35B MoE models aren’t enough and you’re targeting the largest open-weight models (100B to 671B) locally, the “Which LLM on Mac Studio” guide covers Ultra configurations with up to 512 GB.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.