BestLLMfor Your hardware. Your LLM. Your call.
◆ The kits◆ Kits APIOpen data Find my LLM

Which local LLM can my Mac run?

Pick your chip and memory: the tool lists the catalog models that fit, with the quantization to use and the speed to expect.

Your Mac

Loading the catalog…

What runs on each Mac memory size

Unified memory is shared between macOS and the models. By default the GPU can use about two thirds of it up to 32 GB, three quarters above. Suggestion: the largest, most recent model that fits at Q4 with an 8K context.

MemoryUsableSuggestion (Q4)
8 GB≈ 5.4 GBGranite 4.0 H-Tiny 7B-A1B · 8 GB ranking
16 GB≈ 11 GBGemma 4 12B · 16 GB ranking
24 GB≈ 16 GBDevstral Small 2 24B · 24 GB ranking
32 GB≈ 21 GBLaguna XS.2 · 32 GB ranking
36 GB≈ 27 GBQwen 3.6 35B-A3B
48 GB≈ 36 GBQwen 3.6 35B-A3B · 48 GB ranking
64 GB≈ 48 GBNemotron 3 Puzzle 75B-A9B · 64 GB ranking
96 GB≈ 72 GBLaguna S 2.1 · 96 GB ranking
128 GB≈ 96 GBMistral Medium 3.5 128B · 128 GB ranking
192 GB≈ 144 GBMiniMax-M2.7
256 GB≈ 192 GBGLM 5.3 Flash 320B-A18B
512 GB≈ 384 GBNemotron 3 Ultra (BF16)

Speed depends on the chip

For every generated token the Mac reads the whole model from memory, so bandwidth sets the speed ceiling. At equal memory, a Max or Ultra chip is several times faster than a base chip.

ChipBandwidthMemory options
M168 GB/s8, 16 GB
M1 Pro200 GB/s16, 32 GB
M1 Max400 GB/s32, 64 GB
M1 Ultra800 GB/s64, 128 GB
M2100 GB/s8, 16, 24 GB
M2 Pro200 GB/s16, 32 GB
M2 Max400 GB/s32, 64, 96 GB
M2 Ultra800 GB/s64, 128, 192 GB
M3100 GB/s8, 16, 24 GB
M3 Pro150 GB/s18, 36 GB
M3 Max300-400 GB/s36, 48, 64, 96, 128 GB
M3 Ultra819 GB/s96, 256, 512 GB
M4120 GB/s16, 24, 32 GB
M4 Pro273 GB/s24, 48, 64 GB
M4 Max410-546 GB/s36, 48, 64, 128 GB
M5153 GB/s16, 24, 32 GB
M5 Pro307 GB/s24, 48, 64 GB
M5 Max460-614 GB/s36, 48, 64, 96, 128 GB
M5 Ultra1200 GB/s96, 256, 512 GB
M6170 GB/s16, 24, 32 GB

Two values: the version with fewer GPU cores (and less memory) has lower bandwidth. Sources: Apple spec sheets; M5 and M6: our Mac Studio M5 and Mac mini M6 reviews.

Frequently asked questions

Can a Mac with 8 GB run an LLM?

Yes, but only small models: about 5 GB is usable by models, enough for a 3-4B model at Q4 (Qwen 3 4B, Gemma 3 4B). From 16 GB, 7-8B models become comfortable.

How much unified memory can the GPU use?

By default macOS keeps about a third of memory for the system on Macs up to 32 GB, a quarter above that. A 24 GB Mac therefore gives models about 16 GB, a 96 GB Mac about 72 GB. The sysctl iogpu.wired_limit_mb setting raises that limit, as long as macOS keeps enough memory.

Why don't two Macs with the same memory run at the same speed?

Because generation speed mostly depends on memory bandwidth: 120 GB/s on an M4, 273 GB/s on an M4 Pro, 546 GB/s on an M4 Max. For every token, the machine reads the whole model again. An M4 Max is therefore about 4 times faster than an M4 on a model that fits both.

MLX, Ollama or LM Studio on a Mac?

All three work. MLX, Apple's framework, is often the fastest on Apple Silicon; LM Studio offers it with a graphical interface; Ollama is the simplest from the command line and plugs into many tools.

More free tools

All local LLM tools →