Your Mac
Loading the catalog…
What runs on each Mac memory size
Unified memory is shared between macOS and the models. By default the GPU can use about two thirds of it up to 32 GB, three quarters above. Suggestion: the largest, most recent model that fits at Q4 with an 8K context.
| Memory | Usable | Suggestion (Q4) |
|---|---|---|
| 8 GB | ≈ 5.4 GB | Granite 4.0 H-Tiny 7B-A1B · 8 GB ranking |
| 16 GB | ≈ 11 GB | Gemma 4 12B · 16 GB ranking |
| 24 GB | ≈ 16 GB | Devstral Small 2 24B · 24 GB ranking |
| 32 GB | ≈ 21 GB | Laguna XS.2 · 32 GB ranking |
| 36 GB | ≈ 27 GB | Qwen 3.6 35B-A3B |
| 48 GB | ≈ 36 GB | Qwen 3.6 35B-A3B · 48 GB ranking |
| 64 GB | ≈ 48 GB | Nemotron 3 Puzzle 75B-A9B · 64 GB ranking |
| 96 GB | ≈ 72 GB | Laguna S 2.1 · 96 GB ranking |
| 128 GB | ≈ 96 GB | Mistral Medium 3.5 128B · 128 GB ranking |
| 192 GB | ≈ 144 GB | MiniMax-M2.7 |
| 256 GB | ≈ 192 GB | GLM 5.3 Flash 320B-A18B |
| 512 GB | ≈ 384 GB | Nemotron 3 Ultra (BF16) |
Speed depends on the chip
For every generated token the Mac reads the whole model from memory, so bandwidth sets the speed ceiling. At equal memory, a Max or Ultra chip is several times faster than a base chip.
| Chip | Bandwidth | Memory options |
|---|---|---|
| M1 | 68 GB/s | 8, 16 GB |
| M1 Pro | 200 GB/s | 16, 32 GB |
| M1 Max | 400 GB/s | 32, 64 GB |
| M1 Ultra | 800 GB/s | 64, 128 GB |
| M2 | 100 GB/s | 8, 16, 24 GB |
| M2 Pro | 200 GB/s | 16, 32 GB |
| M2 Max | 400 GB/s | 32, 64, 96 GB |
| M2 Ultra | 800 GB/s | 64, 128, 192 GB |
| M3 | 100 GB/s | 8, 16, 24 GB |
| M3 Pro | 150 GB/s | 18, 36 GB |
| M3 Max | 300-400 GB/s | 36, 48, 64, 96, 128 GB |
| M3 Ultra | 819 GB/s | 96, 256, 512 GB |
| M4 | 120 GB/s | 16, 24, 32 GB |
| M4 Pro | 273 GB/s | 24, 48, 64 GB |
| M4 Max | 410-546 GB/s | 36, 48, 64, 128 GB |
| M5 | 153 GB/s | 16, 24, 32 GB |
| M5 Pro | 307 GB/s | 24, 48, 64 GB |
| M5 Max | 460-614 GB/s | 36, 48, 64, 96, 128 GB |
| M5 Ultra | 1200 GB/s | 96, 256, 512 GB |
| M6 | 170 GB/s | 16, 24, 32 GB |
Two values: the version with fewer GPU cores (and less memory) has lower bandwidth. Sources: Apple spec sheets; M5 and M6: our Mac Studio M5 and Mac mini M6 reviews.
Frequently asked questions
Can a Mac with 8 GB run an LLM?
Yes, but only small models: about 5 GB is usable by models, enough for a 3-4B model at Q4 (Qwen 3 4B, Gemma 3 4B). From 16 GB, 7-8B models become comfortable.
How much unified memory can the GPU use?
By default macOS keeps about a third of memory for the system on Macs up to 32 GB, a quarter above that. A 24 GB Mac therefore gives models about 16 GB, a 96 GB Mac about 72 GB. The sysctl iogpu.wired_limit_mb setting raises that limit, as long as macOS keeps enough memory.
Why don't two Macs with the same memory run at the same speed?
Because generation speed mostly depends on memory bandwidth: 120 GB/s on an M4, 273 GB/s on an M4 Pro, 546 GB/s on an M4 Max. For every token, the machine reads the whole model again. An M4 Max is therefore about 4 times faster than an M4 on a model that fits both.
MLX, Ollama or LM Studio on a Mac?
All three work. MLX, Apple's framework, is often the fastest on Apple Silicon; LM Studio offers it with a graphical interface; Ollama is the simplest from the command line and plugs into many tools.
More free tools
- Which LLM for my PC?
- LLM VRAM calculator
- Quantization recommender
- Local LLM budget planner
- LLM cost calculator
- API cost predictor
- Local LLM leaderboard
- Model comparison
- GPU comparison
- Top 3 widget
- Public API
- Chrome extension