Tool · Apple Silicon M1 through M6
Which LLM can run on my Mac ?
Choose your chip and memory: the tool lists the catalog models that fit, along with the quantization to use and the expected speed.
Your Mac
Loading catalog…
What runs based on your Mac's memory
Unified memory is shared between macOS and the models. By default, the GPU can use about two-thirds of it up to 32 GB, and three-quarters above that. Suggestion: use the largest, newest model that fits in Q4 with an 8K context.
| Memory | Usable | Suggestion (Q4) |
|---|---|---|
| 8 GB | ≈ 5.4 GB | LFM2.5 7B · 8 GB tier |
| 16 GB | ≈ 11 GB | Nemotron 3 Super 12B · 16 GB ranking |
| 24 GB | ≈ 16 GB | Devstral Small 2 24B · 24 GB class |
| 32 GB | ≈ 21 GB | Laguna XS.2 · 32 GB ranking |
| 36 GB | ≈ 27 GB | Qwen 3.6 35B-A3B |
| 48 GB | ≈ 36 GB | Qwen 3.6 35B-A3B · 48 GB ranking |
| 64 GB | ≈ 48 GB | Nemotron 3 Puzzle 75B-A9B · 64 GB ranking |
| 96 GB | ≈ 72 GB | Laguna S 2.1 · 96 GB ranking |
| 128 GB | ≈ 96 GB | Mistral Medium 3.5 128B · 128 GB ranking |
| 192 GB | ≈ 144 GB | MiniMax-M2.7 |
| 256 GB | ≈ 192 GB | GLM 5.3 Flash 320B-A18B |
| 512 GB | ≈ 384 GB | Nemotron 3 Ultra (BF16) |
Speed depends on the chip
For every generated word, the Mac rereads the entire model in memory: bandwidth therefore sets the speed ceiling. With the same memory capacity, a Max or Ultra chip is several times faster than a base chip.
| Chip | Bandwidth | Suggested memory |
|---|---|---|
| M1 | 68 GB/s | 8, 16 GB |
| M1 Pro | 200 GB/s | 16, 32 GB |
| M1 Max | 400 GB/s | 32, 64 GB |
| M1 Ultra | 800 GB/s | 64, 128 GB |
| M2 | 100 GB/s | 8, 16, 24 GB |
| M2 Pro | 200 GB/s | 16, 32 GB |
| M2 Max | 400 GB/s | 32, 64, 96 GB |
| M2 Ultra | 800 GB/s | 64, 128, 192 GB |
| M3 | 100 GB/s | 8, 16, 24 GB |
| M3 Pro | 150 GB/s | 18, 36 GB |
| M3 Max | 300 to 400 GB/s | 36, 48, 64, 96, 128 GB |
| M3 Ultra | 819 GB/s | 96, 256, 512 GB |
| M4 | 120 GB/s | 16, 24, 32 GB |
| M4 Pro | 273 GB/s | 24, 48, 64 GB |
| M4 Max | 410 to 546 GB/s | 36, 48, 64, 128 GB |
| M5 | 153 GB/s | 16, 24, 32 GB |
| M5 Pro | 307 GB/s | 24, 48, 64 GB |
| M5 Max | 460 to 614 GB/s | 36, 48, 64, 96, 128 GB |
| M5 Ultra | 1200 GB/s | 96, 256, 512 GB |
| M6 | 170 GB/s | 16, 24, 32 GB |
Two values: the version with fewer GPU cores (and less memory) has lower bandwidth. Sources: Apple technical specifications; M5 and M6: our spec sheets Mac Studio M5 et Mac mini M6.
Frequently asked questions
Can a Mac with 8 GB run an LLM?
Yes, but only small models: about 5 GB is usable by the models, leaving room for a model with 3 to 4 billion parameters in Q4 (Qwen 3 4B, Gemma 3 4B). Starting at 16 GB, 7–8B models become comfortable.
How much unified memory can the GPU use?
By default, macOS reserves about one-third of memory for the system on Macs up to 32 GB, and one-quarter above that. So a 24 GB Mac provides about 16 GB for models, while a 96 GB Mac provides about 72 GB. The sysctl iogpu.wired_limit_mb command can raise this limit while leaving enough memory for macOS.
Why don’t two Macs with the same amount of memory run at the same speed?
Because generation speed depends mainly on memory bandwidth: 120 GB/s on an M4, 273 GB/s on an M4 Pro, and 546 GB/s on an M4 Max. With each word, the machine rereads the entire model. An M4 Max is therefore about 4 times faster than an M4 on a model that fits on both.
MLX, Ollama, or LM Studio on Mac?
All three work. MLX, Apple's framework, is often the fastest on Apple Silicon; LM Studio offers it with a graphical interface; Ollama is the simplest from the command line and integrates with many tools.