M4 Max Mac and local LLM: MLX, performance, models
The MacBook Pro M4 Max (64 or 128 GB of unified memory) has become the most capable portable machine for running a local LLM. The combination of unified memory and Apple's MLX framework makes it possible to load today's largest open-weight models—the 30–35B MoEs at full precision and dense 27B multimodal models—on a laptop, at surprisingly fast speeds. This guide measures what actually works: MLX vs. Ollama, practically usable models, measured tokens/sec, and thermals on battery.
Choosing a machine? Our picks by budget →
Buying alternative for this guide: BOSGAME M5 128GB / 2TB (Ryzen AI Max+ 395).
Why this choice? Our complete guide on BOSGAME M5 128GB / 2TB (Ryzen AI Max+ 395) →
Compare all options by budget, from €800 to €3,500 →
On the go: which laptop for local AI →
Affiliate links — possible commission at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.
#Why the M4 Max changes the game for local LLMs
On a Windows or Linux PC, the GPU's VRAM is the main constraint. A RTX 4090 is capped at 24 GB: running a 30-35B model locally in FP16 requires CPU offloading, which cuts throughput by a factor of five or ten. The M4 Max approaches the problem from the other direction. The GPU and CPU share the same memory: everything in RAM is accessible to the GPU without copying.
In practical terms, a 128 GB M4 Max MacBook Pro can load the best 2026 MoE models — Qwen 3.6 35B-A3B or Qwen 3.8 27B — at full FP16 precision, or several models at once, without breaking a sweat, and keep running them on battery. No portable x86 machine can do that today — you need a desktop with a RTX 5090 or two RTX 3090 to approach the same result, and even then, not on the go.
#Unified memory: why 128 GB ≠ 128 GB of VRAM
You've seen what an M4 Max with MLX can do. The Mac Kit explains when MLX really beats GGUF (ch. 3), how to calculate your memory budget for long contexts (ch. 5), and what performance to expect, including heat and battery life (ch. 6).
- Lifetime online access
- PDF + files
- Lifetime updates
macOS does not let a single process consume all the memory. By default, the GPU can use about 75% of the total RAM. On a 128 GB machine, that gives you ~96 GB available for your models. On a 64 GB machine, you get ~48 GB. That is more than enough for the target models, but you need to know this before calculating.
- Mac M4 Max 64 GB
- About 48 GB for the LLM. The best 2026 MoE models in Q8 fit comfortably—Qwen3-Coder 30B-A3B (≈32 GB), Qwen 3.8 27B (≈30 GB), Qwen 3.6 35B-A3B (≈40 GB), which just fits—and Mistral Small 24B even fits in FP16 (≈48 GB).
- Mac M4 Max 128 GB
- About 96 GB for the LLM. Room to run large MoE models at full FP16 precision (Qwen 3.6 35B-A3B, Qwen 3.8 27B, Granite 4.2 30B), keep multiple models loaded in parallel, or push the context to 256k tokens.
- Recommended quantizations
- Q4_K_M remains the right default. With 128 GB, you can move up to Q5_K_M or Q6 without difficulty for better quality.
#MLX vs Ollama: which one for an M4 Max Mac?
Apple released MLX in December 2023: an array-computing framework designed for Apple Silicon, functionally equivalent to PyTorch but built around unified memory. The mlx-lm module provides a CLI and Python API for running models quantized in the MLX format (similar to GGUF but distinct).
- Ollama (Metal)
- Easier to install, a huge ecosystem, universal GGUF formats, and an OpenAI-compatible API on :11434. Slight overhead, with suboptimal CPU↔GPU memory conversion.
- MLX (mlx-lm)
- Faster in practice on Apple Silicon: +20 to +40% tokens/sec on 30B+ models, native memory management, built-in prompt caching. Smaller ecosystem, with no standardized REST API out of the box.
- Verdict
- To explore and try 10 models: Ollama. To fully leverage the M4 Max and run large 30–35B MoEs at full precision in local production: MLX.
#Install MLX and run a first model
- 01Prepare a Python environmentUse Python 3.10+ (3.12 recommended). On Mac, uv or pyenv is the cleanest option. mlx-lm has no native prerequisites other than Apple Silicon.
- 02Install mlx-lmOne pip command is all it takes. Everything is compiled for arm64 Metal.
- 03Run an MLX modelThe mlx_lm.generate CLI downloads the model from Hugging Face (mlx-community organization) and runs it.
- 04Expose an HTTP API (optional)Since version 0.18, mlx-lm has included an mlx_lm.server command that exposes an OpenAI-compatible API on port 8080.
#Models tested on a MacBook Pro M4 Max 128 GB
All figures below were measured on a 16-inch M4 Max MacBook Pro (40 GPU cores, 128 GB), plugged into power, first cold and then after 5 minutes of warm-up. Short prompt (50 tokens), 500-token generation, batch 1.
- Qwen 3.6 35B-A3B Q8 (MLX)
- ≈40 GB in RAM. 42–50 tokens/s: the MoE activates only 3B parameters, hence the high throughput despite its size. The reliable all-purpose choice on this Mac.
- Qwen 3.8 27B Q8 (MLX)
- ≈30 GB of RAM. 18-22 tokens/s (dense model). 262k context and vision, the closest thing to a local Copilot. Consider setting its reasoning level to low; otherwise, it overthinks.
- Qwen3-Coder 30B-A3B Q8 (MLX)
- ≈32 GB of RAM. 44–52 tokens/s. Code-specialized MoE, 256k context: ideal as a local development assistant.
- GLM 4.7 Flash Q8 (MLX)
- ≈32 GB of RAM (MoE 30B-A3B, MIT license). 40–48 tokens/s. Excellent for agents and tool calls, highly responsive.
- Granite 4.2 30B Q8 (MLX)
- ≈33 GB of RAM. 30–38 tokens/s. Geared toward professional/enterprise use, economical with tokens and stable during long sessions.
- Mistral Small 24B FP16 (MLX)
- ≈48 GB. 18–22 tokens/s. Dense model with excellent French quality, particularly good at document summarization.
#Detailed benchmarks: MLX vs Ollama vs llama.cpp
For the same Qwen 3.8 27B model in Q8, here is the spread observed among the three runtimes on the same M4 Max 128 GB machine:
- 8-bit MLX (mlx-community)
- 20.5 tokens/s on average, prompt eval ~410 tokens/s.
- Ollama Q8_0 (Metal)
- 15.2 tokens/s on average, prompt eval ~300 tokens/s.
- Pure llama.cpp Q8_0
- 15.6 tokens/s on average. Ollama adds little overhead compared with llama.cpp.
#M4 Max vs RTX 4090: who wins at what
The candid comparison: a 128 GB MacBook Pro M4 Max (model replaced by the M5 Max at Apple — look for it used, with variable pricing) versus a desktop equipped with a used RTX 4090 24 GB (variable pricing, GPU only, without the rest of the tower).
- Small models ≤14B Q4 (Qwen 3.5 9B, gpt-oss 20B)
- RTX 4090 wins decisively (80-120 tok/s vs 35-60 on M4 Max). For these sizes, choose the 4090 if you have the choice.
- Dense 24–30B Q4 models (Mistral Small, Qwen 3.8 27B)
- RTX 4090 still ahead (30-40 tok/s vs. 18-22 on M4 Max), but the gap is narrowing.
- Full precision Q8/FP16
- The M4 Max pulls ahead in practice: it runs the best 2026 MoE models (35B-A3B, 27B) in Q8 or FP16 within its RAM, whereas the RTX 4090 has to drop to Q4 or offload to the CPU (throughput divided by five).
- Large MoE (Qwen 3.6 35B-A3B)
- The M4 Max is excellent thanks to its massive RAM: everything fits in memory, even in FP16. RTX 4090 must offload as soon as you exceed Q4.
- Mobility
- No match: the M4 Max gets ~3 hours of continuous inference on battery, while the RTX 4090 goes 0 km.
#Thermals, fans, and battery life
The M4 Max runs relatively cool during inference compared with a NVIDIA GPU. Under sustained load (5 minutes of generation on a large 35B MoE in Q8), the fans ramp up to about 2 800 RPM—audible but far from the noise of a gaming PC. CPU/GPU temperature levels off at around 95 °C without significant throttling on AC power.
- Plugged in (140 W charger)
- Full power, maximum throughput. No limit, moderate fan speed.
- On battery power
- macOS slightly reduces the GPU frequency: ~80% of sector throughput. Qwen 3.6 35B-A3B goes from about 45 to 37 tokens/s.
- Inference battery life
- 16-inch MacBook Pro M4 Max, 100 Wh: ~3 hours of continuous inference on a large 35B MoE, ~5 hours on 7B–14B models, ~8 hours of mixed chat + reading.
- Fan noise
- Audible after 30 seconds of continuous generation, but much quieter than an RTX PC laptop.
#Tips for getting the most out of the M4 Max
- Keep the model loaded
- The first prompt after loading is slow. Use server mode (mlx_lm.server or ollama serve) to keep the model in RAM for the entire session.
- Prefer the MLX format
- Models published by mlx-community on Hugging Face are optimized for Apple Silicon. Look for the "mlx-community/" prefix before anything else.
- Monitor with asitop or Stats
- asitop (the nvtop equivalent for Apple Silicon) shows GPU usage, memory, and instantaneous power draw in real time.
- High-performance mode
- In System Settings → Battery → select “High Power Mode.” The gain is marginal but real for 70B models.
- Disable Spotlight during benchmarks
- Spotlight indexing can consume 5 to 10% of throughput. mdutil -a -i off disables it for the duration of a session.
#Go further
Three ways to explore the topic further:
- Install Ollama on macOS
- If you want simplicity before maximum performance, the guide "Install Ollama on macOS (Apple Silicon)" covers all the Metal tooling on the Ollama side.
- Choose your quantization
- With 128 GB of unified RAM, you have enough headroom to move up to Q5_K_M, Q6, or even FP16—the “Choosing your quantization” guide details the trade-offs.
- Mac Studio Ultra for going further
- If 30–35B MoE models aren’t enough and you’re targeting the largest open-weight models (100B to 671B) locally, the “Which LLM on Mac Studio” guide covers Ultra configurations with up to 512 GB.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.