Mac Studio M5 Max / M5 Ultra: review for local AI
The first desktop that runs a frontier-class LLM entirely on-device. We look at what matters for running LLMs locally — memory, bandwidth, speed, price — and who it is (or is not) the right buy for.
Where to buy
The button names the configuration offered. A mini PC is a complete PC alternative: check usable memory and software support. It does not replace macOS/MLX or CUDA.
As an Amazon Associate, BestLLMfor earns from qualifying purchases, at no extra cost to you. This does not influence our independent recommendations.
Specs
| Memory | M5 Max: 36-128 GB - M5 Ultra: 96-512 GB unified (512 GB ships late October) |
| Bandwidth | M5 Max: 460 GB/s (32-core GPU) / 614 GB/s (40-core GPU) - M5 Ultra: up to 1.2 TB/s (+50% vs previous gen) |
| Compute | M5 Max: 18-core CPU, up to 40-core GPU - M5 Ultra: 36-core CPU, 80-core GPU, 32-core Neural Engine - Neural Accelerators in every GPU core, Thunderbolt 5, PCIe Gen 6 SSD |
| Power | Desktop, quiet cooling |
| Price | from $2,499 (M5 Max 36 GB) - $5,499 (M5 Ultra 96 GB) - $10,000+ fully specced |
| Largest model | 36 GB: 30B class - 64-96 GB: 70B at Q4 - 128 GB: GPT-OSS-120B comfortably - 256 GB: Qwen3-235B (MoE) - 512 GB: ~600B frontier MoE at Q4 (DeepSeek class) - see the memory-tier table |
Who is it for?
macOS developers and creators targeting big local LLMs - from the 30B class (36 GB) up to ~600B frontier MoE at Q4 (512 GB).
Pros and cons
Pros:
- 512 GB of unified memory at 1.2 TB/s: a ~600B-parameter model (Q4) fits ENTIRELY on-device - unheard of outside server racks
- Apple claims up to 4.3x the AI performance of the previous generation (Neural Accelerators per GPU core)
- macOS + MLX: the most optimized local-LLM stack on Apple (LM Studio, Ollama, mlx-lm)
- Quiet, compact, low power for the capacity - no multi-GPU tower fits 512 GB of VRAM
- Thunderbolt 5 and PCIe Gen 6 SSD: much faster weight loading
Cons:
- The 512 GB tier only ships late October 2026 (everything else from September 22)
- Apple pricing: the 512 GB build lands well above $10,000 - serious professional use only
- Below 96 GB, a clearance Mac Studio M4 Max or a 128 GB Strix Halo mini PC gives more memory per dollar
- No CUDA: heavy fine-tuning and some tooling remain simpler on NVIDIA
Which memory tier runs which models?
The rule: memory sets the size of the model you can load, bandwidth sets generation speed. Speeds below are order-of-magnitude estimates for text generation (decode) at Q4 via MLX / llama.cpp - MoE models only read their active parameters per token, hence their high throughput.
| Configuration | Bandwidth | What runs (Q4) | Est. speed |
|---|---|---|---|
| M5 Max 36 GB | 460 GB/s | 30B class: Qwen3-32B, GLM-4.7-Flash | ~25-35 tok/s |
| M5 Max 64 GB | 614 GB/s (40-core GPU) | 70B at Q4 (~40 GB); 30B-A3B MoE very fast | ~13-16 tok/s (70B) - 80+ tok/s (A3B MoE) |
| M5 Max 128 GB | 614 GB/s | 70B at Q8, GPT-OSS-120B (MXFP4), long context | ~40-60 tok/s (GPT-OSS-120B) |
| M5 Ultra 96 GB | 1.2 TB/s | 70B at Q4 with context headroom | ~25-30 tok/s |
| M5 Ultra 256 GB | 1.2 TB/s | Qwen3-235B-A22B at Q4 (~135 GB) | ~25-35 tok/s |
| M5 Ultra 512 GB | 1.2 TB/s | ~600B frontier MoE at Q4 (DeepSeek class, ~380-400 GB) | ~15-20 tok/s |
Prefill (prompt reading) additionally benefits from the M5 GPU Neural Accelerators - historically the weak spot of Macs vs NVIDIA GPUs on long contexts.
M5 Max or M5 Ultra?
The M5 Max ($2,499 at 36 GB, BTO up to 128 GB) covers everything up to 70B and ~120B MoE: the rational pick for demanding dev/copilot/RAG use. Go for the 40-core GPU as soon as you pass 36 GB - bandwidth jumps from 460 to 614 GB/s, a third more speed on every token.
The M5 Ultra ($5,499 at 96 GB) is only justified by its memory: at 96 GB it barely beats a cheaper M5 Max 128 GB. Its real reason to exist is the 256 and 512 GB tiers - the only ones that unlock 235B to ~600B MoE models. No specific use for those models? The Ultra is a bad buy. Have one? It is the only desktop hardware that does it.
Against the alternatives
| Machine | Memory | Bandwidth | Price | Best use |
|---|---|---|---|---|
| Mac Studio M4 Max | 36 GB | 410 GB/s | ~$1,999 (clearance) | The 30B class at the best Apple price |
| Mac Studio M5 Max | 36-128 GB | 460-614 GB/s | from $2,499 | Fast 70B and ~120B MoE, MLX ecosystem |
| Mac Studio M5 Ultra | 96-512 GB | 1.2 TB/s | from $5,499 | 235-600B MoE: no desktop equivalent |
| Strix Halo mini PC | 64-128 GB | 212 GB/s | ~$2,000-3,500 | Best memory per dollar (70B, slower but accessible) |
| NVIDIA DGX Spark | 128 GB | 273 GB/s | ~$4,000-4,700 | CUDA ecosystem, fine-tuning |
| RTX 5090 | 32 GB | 1.79 TB/s | ~$3,000+ | Raw speed up to ~32B, nothing beyond |
In short: raw speed on small models - discrete GPU; max capacity per dollar - Strix Halo; max capacity, period - M5 Ultra.
Pre-order: what to know
- Pre-orders opened August 25, 2026; first deliveries September 22.
- The 512 GB tier only ships late October - order early if that is your target.
- Retail listings currently cover the base M5 Max 36 GB / 512 GB SSD config; higher memory tiers go through Apple BTO.
- The M4 Max 36 GB remains on clearance - see our dedicated review to check whether it is enough for you.
Verdict
The Mac Studio M5 changes the category: until now, running a frontier-class model meant a five-figure multi-GPU server; the M5 Ultra 512 GB does it on a desk, silently. At $2,499 the M5 Max 36 GB is fine but clearance M4 Max stock wins that tier; the M5 makes real sense from 128 GB (614 GB/s), and becomes one-of-a-kind at 256-512 GB. If your workload fits in a 70B, a Strix Halo 128 GB is still about half the price; if you want DeepSeek-class models locally, nothing else comes close.
FAQ
Can the Mac Studio M5 Max / M5 Ultra run a 70B LLM locally?
36 GB: 30B class - 64-96 GB: 70B at Q4 - 128 GB: GPT-OSS-120B comfortably - 256 GB: Qwen3-235B (MoE) - 512 GB: ~600B frontier MoE at Q4 (DeepSeek class) - see the memory-tier table.
How much does the Mac Studio M5 Max / M5 Ultra cost?
Expect from $2,499 (M5 Max 36 GB) - $5,499 (M5 Ultra 96 GB) - $10,000+ fully specced. Prices move fast — check today's price via the buy links.
Who is the Mac Studio M5 Max / M5 Ultra for?
macOS developers and creators targeting big local LLMs - from the 30B class (36 GB) up to ~600B frontier MoE at Q4 (512 GB).
Can the Mac Studio M5 run a DeepSeek-class model locally?
Yes - that is THE headline: with 512 GB of unified memory, a ~600B-parameter frontier MoE quantized to Q4 (~380-400 GB of weights) fits entirely on-device, with generation estimated at 15-20 tok/s thanks to the MoE architecture (only active parameters are read per token).
Should I wait for the M5 or grab a clearance Mac Studio M4 Max?
If your needs fit in 36 GB (30B class), the clearance M4 Max is the better deal: hundreds of dollars less for ~10% less speed. From 64 GB upward, go M5 - higher bandwidth, Neural Accelerators, and memory tiers the M4 Max no longer offers.
Why does unified memory beat a GPU for big models?
A discrete GPU caps at 32 GB of VRAM (RTX 5090): beyond that you offload to system RAM over PCIe and speed collapses. Apple unified memory gives the GPU direct access to the full 96-512 GB: the whole model stays at full bandwidth. Slower than GDDR7 VRAM at equal size, but the only desktop approach that actually FITS giant models.