Which LLM on Mac Studio (M2 / M3 / M4 Ultra, 64–512 GB) ?
In 2026, the Mac Studio (M2 Ultra, M3 Ultra, M4 Max) is the only consumer machine in the world capable of running models with more than 200 billion parameters locally. The 512 GB M3 Ultra (released in early 2025 and replaced in late September 2026 by the M5 Ultra) increased unified memory to 512 GB for about €11,000—enough for Llama 4 Maverick 402B in Q4 or DeepSeek V3 671B in Q4. This guide details each configuration and its use cases.
Choosing a machine? Our picks by budget → · Our spec sheet Mac Studio M5 Max →
Buying alternative for this guide: BOSGAME M5 128GB / 2TB (Ryzen AI Max+ 395).
A mini PC is a complete machine: check the available memory and engine compatibility. It does not replace macOS/MLX or CUDA.
Why this choice? Our complete guide on BOSGAME M5 128GB / 2TB (Ryzen AI Max+ 395) →
Compare all options by budget, from €800 to €3,500 →
On the go: which laptop for local AI →
Affiliate links — possible commission at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.
#Mac Studio in 2026
- Mac Studio M2 Ultra (2023)
- 24 CPU cores + 60 or 76 GPU cores, 800 GB/s, 64/128/192 GB RAM. Still readily available used starting at €2,800.
- Mac Studio M4 Max (2025)
- 16 CPU cores + 40 GPU cores, 546 GB/s, 36–128 GB of RAM. Replaces the M2 Max. No longer sold new by Apple: replaced in late September 2026 by the M5 Max (starting at €2,999).
- Mac Studio M3 Ultra (2025)
- 28 CPU cores + 60 or 80 GPU cores, 800 GB/s, 96/256/512 GB RAM. The monster. Sold for 4499 € (96 GB) to 11 299 € (512 GB) until the M5 Ultra arrived; no longer in the Apple catalog at the end of September 2026 (used or remaining stock).
- Power consumption
- 15 W idle, 70-130 W during inference depending on the configuration. The 512 GB M3 Ultra peaks at ~150 W with 405B.
#1. M2 Ultra vs M3 Ultra vs M4 Max
- M4 Max
- The entry-level Mac Studio. 546 GB/s, 128 GB maximum. Better than an M4 Max MBP: no throttling, lower cost (desktop enclosure).
- M2 Ultra
- First-generation Ultra. 800 GB/s, up to 192 GB. Used price: €2,800–4,500. Still competitive for 70B–123B.
- M3 Ultra
- Latest Ultra (2025). The same 800 GB/s as the M2 Ultra, but with M3 architecture (+20% in practice). 96 to 512 GB of RAM. For anyone looking to go beyond 200B.
#2. How far should you push the RAM?
- 96 GB (base M3 Ultra)
- €4,499 for the M3 Ultra in the catalog, no longer sold new (M5 Ultra 96 GB: €6,599). Runs gpt-oss 120B (MXFP4, 63 GB) at ~50 tok/s and ~30B MoEs (Qwen 3.6 35B-A3B) at ~70 tok/s, with room for a long context. Sweet spot for large MoEs up to ~120B.
- 256 GB (M3 Ultra + 256)
- €7299 listed for the M3 Ultra, no longer sold new. Runs Llama 4 Maverick 402B Q4 (230 GB) at ~34 tok/s and DeepSeek V3 671B Q2 (~240 GB). For anyone who wants 400B+.
- 512 GB (M3 Ultra max)
- €11,299 at the time of the M3 Ultra’s listing, no longer sold new (M5 Ultra 512 GB: late October 2026). DeepSeek V3 671B Q4 (380 GB) works, as does Qwen3-235B-A22B in Q8, with 128k+ context on large MoE models. Research/R&D territory.
- 192 GB (M2 Ultra max)
- Runs Qwen3-235B-A22B Q4 (133 GB) and Maverick 402B in Q3. Functionally equivalent to the 256 GB M3 Ultra for a little less on the used market. A good compromise.
#3. Ultra-exclusive models
| Model | Quant | Required RAM | M2 Ultra 192 | M3 Ultra 256 | M3 Ultra 512 |
|---|---|---|---|---|---|
| gpt-oss 120B | MXFP4 | 63 GB | ✓ 42 t/s | ✓ 50 t/s | ✓ 50 t/s |
| Qwen3-235B-A22B | Q4_K_M | 133 GB | ✓ 26 t/s | ✓ 30 t/s | ✓ 30 t/s |
| Llama 4 Maverick 402B | Q4_K_M | 230 GB | — | ✓ 34 t/s | ✓ 34 t/s |
| Llama 4 Maverick 402B | Q8_0 | 410 GB | — | — | ✓ 22 t/s |
| DeepSeek V3 671B | Q2_K | 240 GB | — | ✓ 20 t/s | ✓ 20 t/s |
| DeepSeek V3 671B | Q4_K_M | 380 GB | — | — | ✓ 16 t/s |
#4. 120B, 235B, and 671B benchmarks
| Model | Q4_K_M | Q5_K_M | Q6_K | Flash Attn 16k |
|---|---|---|---|---|
| gpt-oss 120B | 50 t/s | — | — | 48 t/s |
| Qwen3-235B-A22B | 30 t/s | 28 t/s | 26 t/s | 27 t/s |
| Llama 4 Maverick 402B | 34 t/s | 30 t/s | 28 t/s | 31 t/s |
| DeepSeek V3 671B | 16 t/s | 14 t/s | — | 15 t/s |
#5. LLM workstation setup
#6. DeepSeek V3 671B on M3 Ultra 512 GB
This is the flagship use case for the 512 GB Mac Studio. The model is 380 GB in Q4_K_M, fits in RAM on the 512 GB configuration, and runs at 16k context at ~16 tok/s (MoE, ~37B active parameters out of 671B). Quality is close to the best proprietary models for most tasks.
You know which models a Mac Studio is designed for. The Mac kit provides the reference table by chip and memory capacity (ch. 4), calculates the memory budget for long contexts (ch. 5), and explains how to run Ollama as a background task (ch. 7).
- Lifetime online access
- PDF + files
- Lifetime updates
- Download
- About 380 GB via Hugging Face. Allow 4-8 hours depending on the connection.
- Load time
- ~60 seconds the first time (SSD → RAM). Instant afterward as long as the model remains in memory.
- Inference usage
- ~150 W. Reaches 70-75 °C. Fan is audible but not loud (far quieter than an equivalent PC).
- When it is worth it
- R&D, teams that want a 100% offline frontier model, internal benchmarks, sensitive datasets. For production, the cloud API is more economical.
#Mac Studio vs. PC 2× RTX 5090
- 2026 price
- Mac Studio M3 Ultra 512 GB: ~€11,300 at list price through September 2026; no longer sold new (M5 Ultra 512 GB announced for late October). PC 2× RTX 5090 + platform: ≈ €15,000.
- Useful memory for LLMs
- Mac: 512 GB unified. PC: 2 × 32 GB VRAM = 64 GB. For 400B+ (Maverick 402B, DeepSeek 671B), the PC simply cannot load the model.
- Speed on an ~30B (Qwen 3.6 35B-A3B)
- Dual-5090 PC: ~120 tok/s (CUDA, fits in VRAM). M3 Ultra Mac: ~70 tok/s. PC advantage for what fits in VRAM.
- Speed on DeepSeek 671B / Maverick 402B
- Mac: 16-34 tok/s (MoE). PC: impossible without massive CPU offloading (< 2 tok/s).
- Power consumption
- Mac Studio: 150 W peak. PC with 2× 5090s: 900 W max. Expect the PC electricity bill to be 3× higher.
- Noise
- Mac: slightly audible. PC with 2× 5090s: very loud (GPU fans at full speed).
- Verdict
- The Mac Studio wins on capacity (200B+ models) and efficiency. The PC wins on raw speed for models that fit in VRAM (≤ ~30 GB). Choose based on your target model profile.
#Frequently asked questions
Which Mac Studio for gpt-oss 120B?+
Is the 512 GB Mac Studio M3 Ultra worth €11,000?+
Mac Studio M3 Ultra 256 GB or 512 GB?+
Can you rack-mount a Mac Studio as a server?+
Is the Mac Studio M4 Max better than an MBP M4 Max?+
What is the expected lifespan of a Mac Studio running inference 24/7?+
Prices change quickly: every Monday and Thursday, our tracker records the lowest price for local AI graphics cards, along with the price per GB of VRAM.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.