Home›AI hardware›Mac Studio M5 Max / M5 Ultra

Mac Studio M5 Max / M5 Ultra: local AI test and review

The first consumer machine capable of running a frontier model locally. We look at what matters for running LLMs locally : memory, bandwidth, speed, price—and who it's (or isn't) the right purchase for.

Updated on 09/10/2026

Where to buy it

Mac Studio M5 Max (36 GB / 512 GB)
AmazonPC alternative: GMKtec EVO-X2 64 GB / 1 TB (Ryzen AI Max+ 395) →

The button configuration is the one offered for purchase. On mini PCs, keep some memory available for the system and check the inference engine; macOS/MLX and CUDA are not interchangeable.

Affiliate links — QuelLLM may earn a commission on purchases, at no extra cost to you, which does not influence these recommendations (established independently). As an Amazon Associate, BestLLMfor earns from qualifying purchases.

Technical specifications

Who is it for?

MacOS developers and creators targeting large LLMs on a fixed workstation—from the 30B class (36 GB) up to ~600B frontier MoE models in Q4 (512 GB).

Advantages and limitations

✓ Strengths

  • 512 GB of unified memory at 1.2 TB/s: a ~600-billion-parameter model (Q4) fits ENTIRELY locally — unprecedented outside a server
  • Apple claims up to 4.3× the AI performance of the previous generation (Neural Accelerators per GPU core)
  • macOS + MLX: the most optimized local LLM ecosystem for Apple (LM Studio, Ollama, mlx-lm)
  • Quiet, compact, and low-power for its capacity — no multi-GPU tower can house 512 GB of VRAM
  • Thunderbolt 5 and PCIe Gen 6 SSD: significantly faster weight loading

✗ Limitations

  • The 512 GB tier does not ship until late October 2026 (the rest from 09/22)
  • Price Apple: the 512 GB configuration far exceeds €10,000—reserved for serious professional use
  • Below 96 GB, a 128 GB Strix Halo mini PC (≈ €3,300 to €4,000) offers more memory per euro
  • No CUDA: heavy fine-tuning and some tools remain simpler on the NVIDIA side

Which memory tier for which models?

The rule: the memory sets the size of the model that can be loaded, the bandwidth determines generation speed. Speeds are given as rough orders of magnitude for text generation (decoding), Q4 quantization, via MLX / llama.cpp — MoE models use only their active parameters, hence their high throughput.

ConfigurationBandwidthWhat runs (Q4)Estimated speed
M5 Max 36 GB460 GB/s30B class: Qwen3-32B, GLM-4.7-Flash~25–35 tok/s
M5 Max 64 GB614 GB/s (40-core GPU)70B in Q4 (~40 GB); very fast 30B-A3B MoE~13–16 tok/s (70B) · 80+ tok/s (MoE A3B)
M5 Max 128 GB614 GB/s70B in Q8, GPT-OSS-120B (MXFP4), long context~40-60 tok/s (GPT-OSS-120B)
M5 Ultra 96 GB1.2 TB/s70B in Q4 with ample context headroom~25-30 tok/s
M5 Ultra 256 GB1.2 TB/sQwen3-235B-A22B in Q4 (~135 GB)~25–35 tok/s
M5 Ultra 512 GB1.2 TB/s~600B frontier MoE in Q4 (class DeepSeek, ~380-400 GB)~15-20 tok/s

Prefill (reading the prompt) also benefits from the M5 GPU's Neural Accelerators—this was historically the weak point of Macs compared with NVIDIA GPUs on long contexts.

M5 Max or M5 Ultra?

Le M5 Max (€2,999 for 36 GB, BTO up to 128 GB) covers everything up to 70B and ~120B MoE models: it is the rational choice for demanding dev/copilot/RAG use. Be sure to get the 40-core GPU as soon as you exceed 36 GB: bandwidth increases from 460 to 614 GB/s, delivering one-third more speed for each token.

Le M5 Ultra (€6,599 in 96 GB) is justified only by its memory: at 96 GB it is barely better than a cheaper 128 GB M5 Max. Its real purpose is the tiers 256 and 512 GB — the only ones to run MoE models from 235 to ~600 billion parameters. If you don't have a specific use for these models, the Ultra is a bad purchase; if you do, it's the only desktop hardware that can run them.

Compared with the alternatives

MachineMemoryBandwidthPricingUsing it properly
Mac Studio M5 Max36-128 GB460–614 GB/sstarting at €2,999Fast 70B and ~120B MoE models, MLX ecosystem
Mac Studio M5 Ultra96-512 GB1.2 TB/sfrom €6,599MoE 235–600B: unmatched on desktop
Strix Halo Mini-PC64–128 GB212 GB/s≈ 2 000-4 000 €Best memory-to-euro ratio (70B, slow but accessible)
NVIDIA DGX Spark128 GB273 GB/s≈ 6 100 – 6 600 €CUDA ecosystem required, fine-tuning
RTX 509032 GB1.79 TB/s≈ 5 700 €Raw speed up to ~32B, but no further

In summary: pure speed on small models → dedicated GPU; maximum capacity per dollar → Strix Halo; maximum capacity overall → M5 Ultra.

Availability: what you need to know

Verdict

The Mac Studio M5 changes the category: until now, “running a frontier model” meant an exorbitantly expensive multi-GPU server; the M5 Ultra 512 GB does it on a desktop, silently. At €2,999, the M5 Max 36 GB adequately covers the 30B class; the M5 starts to make real sense at 128 GB (614 GB/s), and becomes unique worldwide at 256-512 GB. If your use case fits within 70B, a Strix Halo 128 GB (≈ €3,300 to €4,000) offers more memory for a price close to an M5 Max, but generates 2 to 3 times more slowly; if you want the DeepSeek class locally, there is simply nothing else.

Frequently asked questions

Can the Mac Studio M5 Max / M5 Ultra run a 70B LLM locally?

36 GB: 30B class · 64–96 GB: 70B in Q4 · 128 GB: GPT-OSS-120B with room to spare · 256 GB: Qwen3-235B (MoE) · 512 GB: frontier MoE ~600B in Q4, DeepSeek class — see the tier table.

How much does the Mac Studio M5 Max / M5 Ultra cost?

Expect €2,999 (M5 Max 36 GB / 512 GB) · €6,599 (M5 Ultra 96 GB / 1 TB) — up to ≈ €20,000 in 512 GB / 16 TB (starting at $2,499 (M5 Max) · $5,499 (M5 Ultra)). Prices change quickly — check the current price through the purchase links.

Who is the Mac Studio M5 Max / M5 Ultra for?

MacOS developers and creators targeting large LLMs on a fixed workstation—from the 30B class (36 GB) up to ~600B frontier MoE models in Q4 (512 GB).

Can you run a DeepSeek-class model locally on the Mac Studio M5?

Yes, this is THE breakthrough: with 512 GB of unified memory, a frontier MoE model with ~600 billion parameters quantized in Q4 (~380-400 GB of weights) fits entirely locally, with an estimated generation speed of 15-20 tok/s thanks to the MoE architecture (only the active parameters are read for each token). See also our DeepSeek V4 Flash guide.

Should you still choose the Mac Studio M4 Max over the M5?

No longer available new from our retailers: M4 Max 36 GB / 512 GB, cleared out at €2,219 over the summer, is no longer in stock at the end of September 2026. The M5 Max takes over at €2,999: higher bandwidth, Neural Accelerators, and memory tiers up to 128 GB that the M4 Max no longer offers.

Why does unified memory beat a GPU for large models?

A dedicated GPU tops out at 32 GB of VRAM (RTX 5090): beyond that, it has to offload to RAM over the PCIe bus, and speed collapses. Unified memory Apple gives the GPU direct access to the full 96-512 GB: the entire model stays at full bandwidth. It's slower than GDDR7 VRAM of the same size, but it's the only desktop approach that can “fit” giant models.

Go further