◆ Mac — Local AI on your Mac, done right — MLX, Ollama, LM Studio on Apple Silicon · $24 · or all kits $49 →

Tool · Apple Silicon M1 through M6

Which LLM can run on my Mac ?

Choose your chip and memory: the tool lists the catalog models that fit, along with the quantization to use and the expected speed.

Your Mac

Loading catalog…

What runs based on your Mac's memory

Unified memory is shared between macOS and the models. By default, the GPU can use about two-thirds of it up to 32 GB, and three-quarters above that. Suggestion: use the largest, newest model that fits in Q4 with an 8K context.

MemoryUsableSuggestion (Q4)
8 GB≈ 5.4 GBLFM2.5 7B · 8 GB tier
16 GB≈ 11 GBNemotron 3 Super 12B · 16 GB ranking
24 GB≈ 16 GBDevstral Small 2 24B · 24 GB class
32 GB≈ 21 GBLaguna XS.2 · 32 GB ranking
36 GB≈ 27 GBQwen 3.6 35B-A3B
48 GB≈ 36 GBQwen 3.6 35B-A3B · 48 GB ranking
64 GB≈ 48 GBNemotron 3 Puzzle 75B-A9B · 64 GB ranking
96 GB≈ 72 GBLaguna S 2.1 · 96 GB ranking
128 GB≈ 96 GBMistral Medium 3.5 128B · 128 GB ranking
192 GB≈ 144 GBMiniMax-M2.7
256 GB≈ 192 GBGLM 5.3 Flash 320B-A18B
512 GB≈ 384 GBNemotron 3 Ultra (BF16)

Speed depends on the chip

For every generated word, the Mac rereads the entire model in memory: bandwidth therefore sets the speed ceiling. With the same memory capacity, a Max or Ultra chip is several times faster than a base chip.

ChipBandwidthSuggested memory
M168 GB/s8, 16 GB
M1 Pro200 GB/s16, 32 GB
M1 Max400 GB/s32, 64 GB
M1 Ultra800 GB/s64, 128 GB
M2100 GB/s8, 16, 24 GB
M2 Pro200 GB/s16, 32 GB
M2 Max400 GB/s32, 64, 96 GB
M2 Ultra800 GB/s64, 128, 192 GB
M3100 GB/s8, 16, 24 GB
M3 Pro150 GB/s18, 36 GB
M3 Max300 to 400 GB/s36, 48, 64, 96, 128 GB
M3 Ultra819 GB/s96, 256, 512 GB
M4120 GB/s16, 24, 32 GB
M4 Pro273 GB/s24, 48, 64 GB
M4 Max410 to 546 GB/s36, 48, 64, 128 GB
M5153 GB/s16, 24, 32 GB
M5 Pro307 GB/s24, 48, 64 GB
M5 Max460 to 614 GB/s36, 48, 64, 96, 128 GB
M5 Ultra1200 GB/s96, 256, 512 GB
M6170 GB/s16, 24, 32 GB

Two values: the version with fewer GPU cores (and less memory) has lower bandwidth. Sources: Apple technical specifications; M5 and M6: our spec sheets Mac Studio M5 et Mac mini M6.

Frequently asked questions

Can a Mac with 8 GB run an LLM?

Yes, but only small models: about 5 GB is usable by the models, leaving room for a model with 3 to 4 billion parameters in Q4 (Qwen 3 4B, Gemma 3 4B). Starting at 16 GB, 7–8B models become comfortable.

How much unified memory can the GPU use?

By default, macOS reserves about one-third of memory for the system on Macs up to 32 GB, and one-quarter above that. So a 24 GB Mac provides about 16 GB for models, while a 96 GB Mac provides about 72 GB. The sysctl iogpu.wired_limit_mb command can raise this limit while leaving enough memory for macOS.

Why don’t two Macs with the same amount of memory run at the same speed?

Because generation speed depends mainly on memory bandwidth: 120 GB/s on an M4, 273 GB/s on an M4 Pro, and 546 GB/s on an M4 Max. With each word, the machine rereads the entire model. An M4 Max is therefore about 4 times faster than an M4 on a model that fits on both.

MLX, Ollama, or LM Studio on Mac?

All three work. MLX, Apple's framework, is often the fastest on Apple Silicon; LM Studio offers it with a graphical interface; Ollama is the simplest from the command line and integrates with many tools.

More free tools