Mac Studio M5 Max / M5 Ultra: review for local AI
The first desktop that runs a frontier-class LLM entirely on-device. We look at what matters for running LLMs locally — memory, bandwidth, speed, price — and who it is (or is not) the right buy for.
Where to buy
As an Amazon Associate, BestLLMfor earns from qualifying purchases, at no extra cost to you. This does not influence our independent recommendations.
Specs
| Memory | M5 Max: 36-128 GB - M5 Ultra: 96-512 GB unified (512 GB ships late October) |
| Bandwidth | M5 Max: 460 GB/s (32-core GPU) / 614 GB/s (40-core GPU) - M5 Ultra: up to 1.2 TB/s (+50% vs previous gen) |
| Compute | M5 Max: 18-core CPU, up to 40-core GPU - M5 Ultra: 36-core CPU, 80-core GPU, 32-core Neural Engine - Neural Accelerators in every GPU core, Thunderbolt 5, PCIe Gen 6 SSD |
| Power | Desktop, quiet cooling |
| Price | from $2,499 (M5 Max 36 GB) - $5,499 (M5 Ultra 96 GB) - $10,000+ fully specced |
| Largest model | 36 GB: 30B class - 64-96 GB: 70B at Q4 - 128 GB: GPT-OSS-120B comfortably - 256 GB: Qwen3-235B (MoE) - 512 GB: ~600B frontier MoE at Q4 (DeepSeek class) - see the memory-tier table |
Who is it for?
macOS developers and creators targeting big local LLMs - from the 30B class (36 GB) up to ~600B frontier MoE at Q4 (512 GB).
Pros and cons
Pros:
- 512 GB of unified memory at 1.2 TB/s: a ~600B-parameter model (Q4) fits ENTIRELY on-device - unheard of outside server racks
- Apple claims up to 4.3x the AI performance of the previous generation (Neural Accelerators per GPU core)
- macOS + MLX: the most optimized local-LLM stack on Apple (LM Studio, Ollama, mlx-lm)
- Quiet, compact, low power for the capacity - no multi-GPU tower fits 512 GB of VRAM
- Thunderbolt 5 and PCIe Gen 6 SSD: much faster weight loading
Cons:
- The 512 GB tier only ships late October 2026 (everything else from September 22)
- Apple pricing: the 512 GB build lands well above $10,000 - serious professional use only
- Below 96 GB, a clearance Mac Studio M4 Max or a 128 GB Strix Halo mini PC gives more memory per dollar
- No CUDA: heavy fine-tuning and some tooling remain simpler on NVIDIA
Which memory tier runs which models?
The rule: memory sets the size of the model you can load, bandwidth sets generation speed. Speeds below are order-of-magnitude estimates for text generation (decode) at Q4 via MLX / llama.cpp - MoE models only read their active parameters per token, hence their high throughput.
| Configuration | Bandwidth | What runs (Q4) | Est. speed |
|---|---|---|---|
| M5 Max 36 GB | 460 GB/s | 30B class: Qwen3-32B, GLM-4.7-Flash | ~25-35 tok/s |
| M5 Max 64 GB | 614 GB/s (40-core GPU) | 70B at Q4 (~40 GB); 30B-A3B MoE very fast | ~13-16 tok/s (70B) - 80+ tok/s (A3B MoE) |
| M5 Max 128 GB | 614 GB/s | 70B at Q8, GPT-OSS-120B (MXFP4), long context | ~40-60 tok/s (GPT-OSS-120B) |
| M5 Ultra 96 GB | 1.2 TB/s | 70B at Q4 with context headroom | ~25-30 tok/s |
| M5 Ultra 256 GB | 1.2 TB/s | Qwen3-235B-A22B at Q4 (~135 GB) | ~25-35 tok/s |
| M5 Ultra 512 GB | 1.2 TB/s | ~600B frontier MoE at Q4 (DeepSeek class, ~380-400 GB) | ~15-20 tok/s |
Prefill (prompt reading) additionally benefits from the M5 GPU Neural Accelerators - historically the weak spot of Macs vs NVIDIA GPUs on long contexts.
M5 Max or M5 Ultra?
The M5 Max ($2,499 at 36 GB, BTO up to 128 GB) covers everything up to 70B and ~120B MoE: the rational pick for demanding dev/copilot/RAG use. Go for the 40-core GPU as soon as you pass 36 GB - bandwidth jumps from 460 to 614 GB/s, a third more speed on every token.
The M5 Ultra ($5,499 at 96 GB) is only justified by its memory: at 96 GB it barely beats a cheaper M5 Max 128 GB. Its real reason to exist is the 256 and 512 GB tiers - the only ones that unlock 235B to ~600B MoE models. No specific use for those models? The Ultra is a bad buy. Have one? It is the only desktop hardware that does it.
Against the alternatives
| Machine | Memory | Bandwidth | Price | Best use |
|---|---|---|---|---|
| Mac Studio M4 Max | 36 GB | 410 GB/s | ~$1,999 (clearance) | The 30B class at the best Apple price |
| Mac Studio M5 Max | 36-128 GB | 460-614 GB/s | from $2,499 | Fast 70B and ~120B MoE, MLX ecosystem |
| Mac Studio M5 Ultra | 96-512 GB | 1.2 TB/s | from $5,499 | 235-600B MoE: no desktop equivalent |
| Strix Halo mini PC | 64-128 GB | 212 GB/s | ~$2,000-3,500 | Best memory per dollar (70B, slower but accessible) |
| NVIDIA DGX Spark | 128 GB | 273 GB/s | ~$4,000-4,700 | CUDA ecosystem, fine-tuning |
| RTX 5090 | 32 GB | 1.79 TB/s | ~$3,000+ | Raw speed up to ~32B, nothing beyond |
In short: raw speed on small models - discrete GPU; max capacity per dollar - Strix Halo; max capacity, period - M5 Ultra.
Pre-order: what to know
- Pre-orders opened August 25, 2026; first deliveries September 22.
- The 512 GB tier only ships late October - order early if that is your target.
- Retail listings currently cover the base M5 Max 36 GB / 512 GB SSD config; higher memory tiers go through Apple BTO.
- The M4 Max 36 GB remains on clearance - see our dedicated review to check whether it is enough for you.
Verdict
The Mac Studio M5 changes the category: until now, running a frontier-class model meant a five-figure multi-GPU server; the M5 Ultra 512 GB does it on a desk, silently. At $2,499 the M5 Max 36 GB is fine but clearance M4 Max stock wins that tier; the M5 makes real sense from 128 GB (614 GB/s), and becomes one-of-a-kind at 256-512 GB. If your workload fits in a 70B, a Strix Halo 128 GB is still about half the price; if you want DeepSeek-class models locally, nothing else comes close.
FAQ
Can the Mac Studio M5 Max / M5 Ultra run a 70B LLM locally?
36 GB: 30B class - 64-96 GB: 70B at Q4 - 128 GB: GPT-OSS-120B comfortably - 256 GB: Qwen3-235B (MoE) - 512 GB: ~600B frontier MoE at Q4 (DeepSeek class) - see the memory-tier table.
How much does the Mac Studio M5 Max / M5 Ultra cost?
Expect from $2,499 (M5 Max 36 GB) - $5,499 (M5 Ultra 96 GB) - $10,000+ fully specced. Prices move fast — check today's price via the buy links.
Who is the Mac Studio M5 Max / M5 Ultra for?
macOS developers and creators targeting big local LLMs - from the 30B class (36 GB) up to ~600B frontier MoE at Q4 (512 GB).
Can the Mac Studio M5 run a DeepSeek-class model locally?
Yes - that is THE headline: with 512 GB of unified memory, a ~600B-parameter frontier MoE quantized to Q4 (~380-400 GB of weights) fits entirely on-device, with generation estimated at 15-20 tok/s thanks to the MoE architecture (only active parameters are read per token).
Should I wait for the M5 or grab a clearance Mac Studio M4 Max?
If your needs fit in 36 GB (30B class), the clearance M4 Max is the better deal: hundreds of dollars less for ~10% less speed. From 64 GB upward, go M5 - higher bandwidth, Neural Accelerators, and memory tiers the M4 Max no longer offers.
Why does unified memory beat a GPU for big models?
A discrete GPU caps at 32 GB of VRAM (RTX 5090): beyond that you offload to system RAM over PCIe and speed collapses. Apple unified memory gives the GPU direct access to the full 96-512 GB: the whole model stays at full bandwidth. Slower than GDDR7 VRAM at equal size, but the only desktop approach that actually FITS giant models.