Advanced 13 minMac Studio

Which LLM on Mac Studio (M2 / M3 / M4 Ultra, 64–512 GB) ?

In 2026, the Mac Studio (M2 Ultra, M3 Ultra, M4 Max) is the only consumer machine in the world capable of running models with more than 200 billion parameters locally. The 512 GB M3 Ultra (released in early 2025 and replaced in late September 2026 by the M5 Ultra) increased unified memory to 512 GB for about €11,000—enough for Llama 4 Maverick 402B in Q4 or DeepSeek V3 671B in Q4. This guide details each configuration and its use cases.

Choosing a machine? Our picks by budget → · Our spec sheet Mac Studio M5 Max →

By Mohamed Meguedmi·Update 2026-08-27·Tested on macOS 14+
Recommended hardware

Buying alternative for this guide: BOSGAME M5 128GB / 2TB (Ryzen AI Max+ 395).

A mini PC is a complete machine: check the available memory and engine compatibility. It does not replace macOS/MLX or CUDA.

Why this choice? Our complete guide on BOSGAME M5 128GB / 2TB (Ryzen AI Max+ 395) →

Compare all options by budget, from €800 to €3,500 →

On the go: which laptop for local AI →

Affiliate links — possible commission at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.

#Mac Studio in 2026

Mac Studio M2 Ultra (2023)
24 CPU cores + 60 or 76 GPU cores, 800 GB/s, 64/128/192 GB RAM. Still readily available used starting at €2,800.
Mac Studio M4 Max (2025)
16 CPU cores + 40 GPU cores, 546 GB/s, 36–128 GB of RAM. Replaces the M2 Max. No longer sold new by Apple: replaced in late September 2026 by the M5 Max (starting at €2,999).
Mac Studio M3 Ultra (2025)
28 CPU cores + 60 or 80 GPU cores, 800 GB/s, 96/256/512 GB RAM. The monster. Sold for 4499 € (96 GB) to 11 299 € (512 GB) until the M5 Ultra arrived; no longer in the Apple catalog at the end of September 2026 (used or remaining stock).
Power consumption
15 W idle, 70-130 W during inference depending on the configuration. The 512 GB M3 Ultra peaks at ~150 W with 405B.
→
The memory card is the real product
The Mac Studio M3 Ultra is more than just a fast computer: it is the only “consumer” machine offering 512 GB of RAM accessible to the GPU. No PC, even at €20,000, offers that natively. For anyone wanting to run 400B+ locally, there is no alternative.

#1. M2 Ultra vs M3 Ultra vs M4 Max

M4 Max
The entry-level Mac Studio. 546 GB/s, 128 GB maximum. Better than an M4 Max MBP: no throttling, lower cost (desktop enclosure).
M2 Ultra
First-generation Ultra. 800 GB/s, up to 192 GB. Used price: €2,800–4,500. Still competitive for 70B–123B.
M3 Ultra
Latest Ultra (2025). The same 800 GB/s as the M2 Ultra, but with M3 architecture (+20% in practice). 96 to 512 GB of RAM. For anyone looking to go beyond 200B.
!
Why no M4 Ultra in 2026?
Apple did not release an M4 Ultra. The M4 lineup ends with the Max. For an Ultra, you had to choose the M3 Ultra (March 2025), before the M5 Ultra arrived, listed by Apple at €6,599 as of late September 2026.

#2. How far should you push the RAM?

96 GB (base M3 Ultra)
€4,499 for the M3 Ultra in the catalog, no longer sold new (M5 Ultra 96 GB: €6,599). Runs gpt-oss 120B (MXFP4, 63 GB) at ~50 tok/s and ~30B MoEs (Qwen 3.6 35B-A3B) at ~70 tok/s, with room for a long context. Sweet spot for large MoEs up to ~120B.
256 GB (M3 Ultra + 256)
€7299 listed for the M3 Ultra, no longer sold new. Runs Llama 4 Maverick 402B Q4 (230 GB) at ~34 tok/s and DeepSeek V3 671B Q2 (~240 GB). For anyone who wants 400B+.
512 GB (M3 Ultra max)
€11,299 at the time of the M3 Ultra’s listing, no longer sold new (M5 Ultra 512 GB: late October 2026). DeepSeek V3 671B Q4 (380 GB) works, as does Qwen3-235B-A22B in Q8, with 128k+ context on large MoE models. Research/R&D territory.
192 GB (M2 Ultra max)
Runs Qwen3-235B-A22B Q4 (133 GB) and Maverick 402B in Q3. Functionally equivalent to the 256 GB M3 Ultra for a little less on the used market. A good compromise.

#3. Ultra-exclusive models

Models that only a Mac Studio Ultra can load comfortably
ModelQuantRequired RAMM2 Ultra 192M3 Ultra 256M3 Ultra 512
gpt-oss 120BMXFP463 GB✓ 42 t/s✓ 50 t/s✓ 50 t/s
Qwen3-235B-A22BQ4_K_M133 GB✓ 26 t/s✓ 30 t/s✓ 30 t/s
Llama 4 Maverick 402BQ4_K_M230 GB—✓ 34 t/s✓ 34 t/s
Llama 4 Maverick 402BQ8_0410 GB——✓ 22 t/s
DeepSeek V3 671BQ2_K240 GB—✓ 20 t/s✓ 20 t/s
DeepSeek V3 671BQ4_K_M380 GB——✓ 16 t/s

#4. 120B, 235B, and 671B benchmarks

Orders of magnitude, Mac Studio M3 Ultra 80-core GPU 512 GB (llama.cpp Metal, 16k context)
ModelQ4_K_MQ5_K_MQ6_KFlash Attn 16k
gpt-oss 120B50 t/s——48 t/s
Qwen3-235B-A22B30 t/s28 t/s26 t/s27 t/s
Llama 4 Maverick 402B34 t/s30 t/s28 t/s31 t/s
DeepSeek V3 671B16 t/s14 t/s—15 t/s

#5. LLM workstation setup

Mac Studio M3 Ultra 512 GB setup
# Ollama + optimisations max
brew install --cask ollama

launchctl setenv OLLAMA_FLASH_ATTENTION 1
launchctl setenv OLLAMA_KV_CACHE_TYPE q8_0
launchctl setenv OLLAMA_NUM_PARALLEL 4
launchctl setenv OLLAMA_HOST "0.0.0.0:11434"

# Relever mémoire GPU à 480 Go (sur 512 Go)
sudo sysctl iogpu.wired_limit_mb=491520
echo "iogpu.wired_limit_mb=491520" | sudo tee -a /etc/sysctl.conf

brew services start ollama

# Pour DeepSeek V3 671B : passer par llama.cpp maison
git clone https://github.com/ggerganov/llama.cpp && cd llama.cpp
make GGML_METAL=1 LLAMA_METAL_EMBED_LIBRARY=1

# Lancement
./build/bin/llama-server \
  -m ./models/DeepSeek-V3-671B-Q4_K_M.gguf \
  -ngl 99 -fa -c 16384 \
  -ctk q8_0 -ctv q8_0 \
  --host 0.0.0.0 --port 8080

#6. DeepSeek V3 671B on M3 Ultra 512 GB

This is the flagship use case for the 512 GB Mac Studio. The model is 380 GB in Q4_K_M, fits in RAM on the 512 GB configuration, and runs at 16k context at ~16 tok/s (MoE, ~37B active parameters out of 671B). Quality is close to the best proprietary models for most tasks.

The Mac Kit

You know which models a Mac Studio is designed for. The Mac kit provides the reference table by chip and memory capacity (ch. 4), calculates the memory budget for long contexts (ch. 5), and explains how to run Ollama as a background task (ch. 7).

  • Lifetime online access
  • PDF + files
  • Lifetime updates
Download
About 380 GB via Hugging Face. Allow 4-8 hours depending on the connection.
Load time
~60 seconds the first time (SSD → RAM). Instant afterward as long as the model remains in memory.
Inference usage
~150 W. Reaches 70-75 °C. Fan is audible but not loud (far quieter than an equivalent PC).
When it is worth it
R&D, teams that want a 100% offline frontier model, internal benchmarks, sensitive datasets. For production, the cloud API is more economical.

#Mac Studio vs. PC 2× RTX 5090

2026 price
Mac Studio M3 Ultra 512 GB: ~€11,300 at list price through September 2026; no longer sold new (M5 Ultra 512 GB announced for late October). PC 2× RTX 5090 + platform: ≈ €15,000.
Useful memory for LLMs
Mac: 512 GB unified. PC: 2 × 32 GB VRAM = 64 GB. For 400B+ (Maverick 402B, DeepSeek 671B), the PC simply cannot load the model.
Speed on an ~30B (Qwen 3.6 35B-A3B)
Dual-5090 PC: ~120 tok/s (CUDA, fits in VRAM). M3 Ultra Mac: ~70 tok/s. PC advantage for what fits in VRAM.
Speed on DeepSeek 671B / Maverick 402B
Mac: 16-34 tok/s (MoE). PC: impossible without massive CPU offloading (< 2 tok/s).
Power consumption
Mac Studio: 150 W peak. PC with 2× 5090s: 900 W max. Expect the PC electricity bill to be 3× higher.
Noise
Mac: slightly audible. PC with 2× 5090s: very loud (GPU fans at full speed).
Verdict
The Mac Studio wins on capacity (200B+ models) and efficiency. The PC wins on raw speed for models that fit in VRAM (≤ ~30 GB). Choose based on your target model profile.

#Frequently asked questions

Which Mac Studio for gpt-oss 120B?+
The model is 63 GB in MXFP4, so you need at least a 96 GB Mac Studio: a used M3 Ultra with 96 GB, or a new M5 Max with 96 GB or M5 Ultra with 96 GB (€6,599)—an M2 Ultra with 64 GB does not leave enough headroom for context. Expect ~50 tok/s on the M3 Ultra, with room for a long 128k context.
Is the 512 GB Mac Studio M3 Ultra worth €11,000?+
For a researcher, an R&D team, or a company handling sensitive data that must not be sent to the cloud: absolutely. No competitor runs Llama 4 Maverick 402B or DeepSeek 671B locally at this price. For an individual, it is probably overkill — a 128 GB Mac Studio (used M4 Max or new M5 Max) is enough for 95% of use cases.
Mac Studio M3 Ultra 256 GB or 512 GB?+
At the catalog prices for the M3 Ultra (no longer sold new): 256 GB (€7,299) if your maximum target is Llama 4 Maverick 402B Q4 (230 GB) or DeepSeek V3 Q2. 512 GB (€11,299) if you want Maverick in Q8 or DeepSeek V3 Q4 (380 GB). The €4,000 difference is justified only if you are truly pushing beyond 400B.
Can you rack-mount a Mac Studio as a server?+
Yes, 19-inch rack solutions exist that accommodate 2 to 4 Mac mini/Studio units. Brands: Sonnet, MacStadium. Useful if you host an LLM cluster for a team.
Is the Mac Studio M4 Max better than an MBP M4 Max?+
Same chip, same performance. Mac Studio wins on price (no display, keyboard, or battery to pay for) and on the complete absence of throttling—sustained 24/7 inference is possible. MBP wins on portability. For stationary use, choose Studio without hesitation.
What is the expected lifespan of a Mac Studio running inference 24/7?+
Apple does not publish official figures, but community reports indicate 5–8 years without hardware issues under continuous inference. The component most at risk is the SSD (write endurance), but LLMs are read and rarely written—minimal SSD wear.

Prices change quickly: every Monday and Thursday, our tracker records the lowest price for local AI graphics cards, along with the price per GB of VRAM.

Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.