Which LLM on a Mac mini M4 / M4 Pro (16–64 GB) ?
A Mac mini M4 runs a local LLM well starting at 16 GB: models with 7 to 12 billion parameters, at best around twenty tokens/s on a 7B in Q4_0, because its 120 GB/s of bandwidth limits speed. To target 14 to 35 billion parameters, you need an M4 Pro (273 GB/s, 24 to 64 GB). Apple replaced this lineup in August 2026 with the Mac mini M6 and M5 Pro.
This guide puts the figures from Apple, llama.cpp, and the Ollama and LM Studio documentation to work on one decision: which chip, how much memory, which model, and what speed to expect. It distinguishes what Apple claims, what the community has measured, and what is merely a calculation, and reviews the M6 and M5 Pro lineup released in September 2026.
Choosing a machine? Our picks by budget →
Buying alternative for this guide: GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395).
A mini PC is a complete machine: check the available memory and engine compatibility. It does not replace macOS/MLX or CUDA.
Why this choice? Our complete guide on GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395) →
Compare all options by budget, from €800 to €3,500 →
On the go: which laptop for local AI →
Affiliate links — possible commission at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.
#Mac mini M4 and local LLMs: what the machine can do
A Mac mini makes a good local LLM server because its unified memory is shared between the CPU and GPU, with no separate VRAM. Two figures determine what you can do with it. Capacity (16 to 64 GB depending on the chip) determines the model size that can load: 7 to 12 billion parameters on 16 GB, and 24 to 35 billion on 32 to 48 GB. Bandwidth determines speed: 120 GB/s for the M4 chip and 273 GB/s for the M4 Pro, according to Apple. An LLM rereads its weights for every generated token, so maximum throughput is roughly the bandwidth divided by the weight size. For a 7B model in Q4_0, llama.cpp's public table reports 24.1 tokens/s for an M4 and about 50 for an M4 Pro. Apple introduced Mac mini M6 and M5 Pro models in August 2026 that replace this lineup; the M4 was still available from some retailers in late August.
- Mac mini M4
- 10-core CPU, 10-core GPU, 120 GB/s, 16 GB unified memory (24 or 32 GB optional), Thunderbolt 4.
- Mac mini M4 Pro
- 12-core CPU, 16-core GPU (14 and 20 optional), 273 GB/s, 24 GB of memory (48 or 64 GB optional), Thunderbolt 5.
- Format and noise
- 12.7 × 12.7 × 5.0 cm, 0.67 kg (M4) or 0.73 kg (M4 Pro); 5 dBA at idle, according to Apple.
- Advertised consumption
- M4: 4 W at idle, 65 W at maximum. M4 Pro: 5 W and 140 W. Apple does not publish an inference figure.
#Mac mini M4 or M4 Pro: bandwidth makes the difference
You know which model fits on your M4 Mac mini. The Mac kit teaches you how to get the most out of it: push the GPU memory limit further (Ch. 2), choose between MLX and GGUF (Ch. 3), and turn your Mac mini into an AI server for the whole house (Ch. 12).
- Lifetime online access
- PDF + files
- Lifetime updates
Between the two chips, the real difference is not maximum memory (32 GB versus 64 GB) but bandwidth, multiplied by 2.3 (273 ÷ 120). In the llama.cpp table, the measured ratio is 2.1 for generation (50.7 ÷ 24.1 in Q4_0) and 2.0 for prompt processing (440 ÷ 221).
| Criterion | Mac mini M4 | Mac mini M4 Pro |
|---|---|---|
| Memory bandwidth (Apple) | 120 GB/s | 273 GB/s |
| Unified memory (Apple) | 16 GB, with 24 or 32 GB options | 24 GB, with a 48 or 64 GB option |
| Generation, Llama 2 7B in Q4_0 (llama.cpp) | 24.1 tokens/s | 50.7 tokens/s (20-core GPU) |
| Generation, same model in Q8_0 | 13.5 tokens/s | 30.7 tokens/s |
| Prompt processing, Q4_0 | 221 tokens/s | 440 tokens/s |
| Thunderbolt ports (Apple) | Thunderbolt 4, 40 Gb/s | Thunderbolt 5, 120 Gb/s |
- The M4 is enough if
- you are on your own, target models with up to 12 billion parameters (Qwen 3.5 9B, Gemma 4 12B) and accept around twenty tokens/s at best. Choose 24 GB to keep an embeddings model loaded as well.
- The M4 Pro is the clear choice if
- you are targeting models with 14 to 27 billion parameters, a 30- to 35-billion-parameter MoE (48 GB minimum), a coding agent, or multiple users. Beyond 64 GB, you need a Mac Studio.
#VRAM on Mac mini: unified memory and its GPU ceiling
A Mac mini has no dedicated VRAM, but the GPU cannot use all the memory. According to llama.cpp users, the default ceiling is two-thirds to three-quarters of the RAM; at launch, llama.cpp displays the recommendedMaxWorkingSetSize line, which gives the actual value. The iogpu.wired_limit_mb setting defaults to 0, meaning “system default policy,” not “unlimited.”
| Mac mini memory | Estimated GPU ceiling | Remains for macOS |
|---|---|---|
| 16 GB | 10.7 to 12 GB | 4 to 5 GB |
| 24 GB | 16 to 18 GB | 6 to 8 GB |
| 32 GB | 21 to 24 GB | 8 to 11 GB |
| 48 GB | 32 to 36 GB | 12 to 16 GB |
| 64 GB | 43 to 48 GB | 16 to 21 GB |
The ceiling must hold the weights, the KV cache (the context memory, which grows with the conversation), and compute buffers. To load a model that is close to the limit, raise the ceiling: MLX recommends a value higher than the model size in megabytes but lower than the machine’s memory. Leave several gigabytes for macOS, or speed will collapse. A manual sysctl setting may not survive a reboot.
#Which model for which memory: what fits in a Mac mini
The table compares the file sizes published by Ollama with a cap of two-thirds of RAM (a conservative assumption). “Yes”: the files use less than 70% of the cap, leaving room for the KV cache. “Barely”: they fit, with a short context. “No”: they exceed it. These are calculations, not measurements.
| Model | Weights (Ollama) | M4 16 GB | M4 24 GB | M4 32 GB | M4 Pro 48 GB | M4 Pro 64 GB |
|---|---|---|---|---|---|---|
| Qwen 3.5 9B | 6.6 GB | yes | yes | yes | yes | yes |
| Gemma 4 12B | 7.6 GB | just | yes | yes | yes | yes |
| gpt-oss 20B | 14 GB | no | just | yes | yes | yes |
| Mistral Small 24B | 14 GB | no | just | yes | yes | yes |
| Qwen 3.8 27B | 18 GB | no | no | just | yes | yes |
| Qwen3-Coder 30B-A3B | 19 GB | no | no | just | yes | yes |
| Qwen 3.6 35B-A3B (MoE) | 23 GB | no | no | no | just | yes |
| Qwen3-Coder 30B-A3B in Q8_0 | 32 GB | no | no | no | just | just |
#Mac mini M4 16 GB: models to prioritize
On 16 GB, the ceiling is around 11 GB. Qwen 3.5 9B (6.6 GB at Ollama) leaves about 4 GB for the KV cache and buffers. Gemma 4 12B (7.6 GB) works with a short context. A 14 GB model such as Mistral Small 24B overflows: Ollama splits it between the GPU and CPU, and speed drops. The ollama ps command displays the split; aim for 100% GPU.
#Mac mini speed: theoretical ceiling and published measurements
Generation throughput is limited by bandwidth: tokens per second ≈ bandwidth (GB/s) ÷ weights read per token (GB). This calculation gives an upper bound, never a measurement. For a dense 14 GB model, 120 ÷ 14 gives 8.6 tokens/s at best on an M4, and 273 ÷ 14 gives 19.5 on an M4 Pro.
| Model | Weights | M4 (120 GB/s) | M4 Pro (273 GB/s) |
|---|---|---|---|
| Qwen 3.5 9B | 6.6 GB | 18 | 41 |
| Gemma 4 12B | 7.6 GB | 16 | 36 |
| Mistral Small 24B | 14 GB | 8,6 | 19,5 |
The published measurements remain below this ceiling. Here are the figures from the table that the llama.cpp community maintains for the Apple chips, with Llama 2 7B. These are not our measurements, and they cover an older model in Q4_0 and Q8_0, but they show the performance ratio between the chips.
| Chip | Bandwidth | Q4_0 generation | Q8_0 generation | Q4_0 prompt |
|---|---|---|---|---|
| M4 (10-core GPU) | 120 GB/s | 24,1 | 13,5 | 221 |
| M4 Pro (16-core GPU) | 273 GB/s | 49,6 | 30,5 | 364 |
| M4 Pro (20-core GPU) | 273 GB/s | 50,7 | 30,7 | 440 |
| M5 Pro (20-core GPU) | 307 GB/s | 66,3 | 38,9 | 1 621 |
Moving from Q8_0 to Q4_0 speeds up generation by a factor of 1.8 on M4 and 1.7 on M4 Pro: file size matters more than computation. GPU cores matter mainly for prompt processing (364 versus 440 tokens/s on M4 Pro). The table contains no M6 measurements as of the date consulted.
#Mac mini M6 and M5 Pro: what the new lineup changes for an LLM
Apple announces up to 4.8 times faster prompt processing on the Mac mini M6 in LM Studio than on the M4, and up to 4 times faster on the M5 Pro than on the M4 Pro. These gains apply to prompt processing, which benefits from the neural accelerators added to each GPU core. Generation remains governed by memory bandwidth and improves much less: from 12% (307 versus 273 GB/s) to 42% (170 versus 120 GB/s).
| Chip | Bandwidth | Unified memory | Idle / maximum |
|---|---|---|---|
| M4 | 120 GB/s | 16, 24, or 32 GB | 4 W / 65 W |
| M6 | 153 GB/s (16 GB), 170 GB/s (24 GB and up) | 16, 24, or 32 GB | 4 W / 70 W |
| M4 Pro | 273 GB/s | 24, 48, or 64 GB | 5 W / 140 W |
| M5 Pro | 307 GB/s | 24, 48, or 64 GB | 6 W / 145 W |
A base M6 at 153 GB/s has only 28% more bandwidth than the M4: the theoretical ceiling for a 9B rises from about 18 to 23 tokens/s. The new lineup is mainly beneficial for long-context use cases: RAG that injects long documents, or a coding agent that rereads a repository.
- Nine, 30- to 35-billion-parameter model
- An M5 Pro with at least 48 GB, or 64 GB to work comfortably with a 35-billion-parameter MoE: the M6 tops out at 32 GB.
- M4 still in stock
- A reasonable choice for 7 to 12 billion parameters if the price gap with the M6 is significant. This stock is limited.
#Install an LLM server on a Mac mini with Ollama
Ollama on macOS works like an application, with its settings configured through environment variables defined with launchctl. By default, Ollama listens only on 127.0.0.1, port 11434: without changes, no other device can use it.
- 01Install the Ollama appDownload Ollama from ollama.com and place the app in Applications. On first launch, it offers to create the ollama command link.
- 02Open the port to the local networkSet OLLAMA_HOST with launchctl (command below), then restart the application. The address 0.0.0.0 opens the API to the entire network, without authentication: limit it to a trusted network.
- 03Prevent sleepIn System Settings, under Energy, enable the option that prevents automatic sleep when the display is off; without it, the service may stop.
- 04Download a suitable modelIn the table, choose a model marked “yes” for your memory, such as qwen3.5:9b on 16 GB.
- 05Check from another deviceSend a request to the Mac mini’s address (.local name or IP), then run ollama ps: the Processor column should show 100% GPU.
Two settings often copied from elsewhere are useless here: Ollama's documentation says Flash Attention is enabled automatically when the hardware allows it, and OLLAMA_ORIGINS concerns browser extensions. By contrast, quantizing the KV cache (OLLAMA_KV_CACHE_TYPE=q8_0) cuts its size roughly in half for a very small loss, which matters on 16 GB when the context grows.
#LM Studio: the alternative to Ollama
LM Studio requires an Apple Silicon chip, macOS 14 or later, and 16 GB of recommended RAM. It loads GGUF models like MLX, starting with version 0.3.4. Without a display, its llmster daemon starts with lms daemon up. The “Serve on Local Network” option exposes the API to the network; token authentication is not enabled by default. Ollama’s MLX engine, presented in preview in March 2026, required more than 32 GB: an M4 16 GB did not benefit from it.
#What a Mac mini can serve: chat, RAG, agents, cluster
- Self-hosted chat server
- The Ollama and LM Studio APIs accept OpenAI-format clients (Open WebUI, Zed, scripts). Placed near the router, the Mac mini serves the whole house.
- Document RAG
- You need to load both the generation model and an embeddings model. nomic-embed-text weighs 274 MB: on 16 GB, it fits alongside Qwen 3.5 9B.
- Code agent
- An agent sends long prompts at every step: prompt reading matters more than generation, and that's where M6 and M5 Pro gain the most.
#How many concurrent users?
By default, Ollama handles one request at a time per model (OLLAMA_NUM_PARALLEL is 1). Each parallel request increases the allocated context: according to the documentation, 4 requests with a context of 2,000 tokens allocate 8,000 tokens. Cache memory = cache for one conversation × parallel requests. Each user gets a slower response; no reliable measurement is available here for a specific number of users.
#Link multiple Mac minis into a cluster
Only the M4 Pro offers Thunderbolt 5 (120 Gb/s); the M4 is limited to Thunderbolt 4 (40 Gb/s). The exo project supports RDMA over Thunderbolt 5, which would reduce latency between machines by 99%, and says it speeds up models as devices are added, on macOS 26.2 or later. Its benchmarks cover Mac Studio, not Mac mini. A cluster is primarily used to load a model that is too large for one machine.
- Share Ollama on your local network (family, team)
- Securing your Ollama server: authentication and reverse proxy
- Exo: turn multiple machines into a home LLM cluster
- Optimize your Mac Apple Silicon: GPU ceiling, Flash Attention, KV cache
- Hardware profile: Mac mini M6
- Hardware profile: Mac mini M5 Pro
#Mac mini M4 or MacBook Air M4: same chip, different use case
The 13-inch M4 MacBook Air uses the same chip, with the same memory bandwidth (120 GB/s) and the same 16, 24, or 32 GB options: the speed ceiling is identical. What changes is the chassis. The previous version of this guide estimated a 20% slowdown for the Air after ten minutes; no source confirms this, so it has been removed. For a service running continuously, the Mac mini is the natural choice; for a mobile workstation (1.24 kg, including the display and battery), the MacBook.
- Source: Apple, Mac mini (2024) technical specifications
- Source: Apple, announcement of the M6 and M5 Pro Mac mini (August 2026)
- Source: llama.cpp, performance on Apple M-series chips
- Source: Ollama documentation, FAQ (environment variables, KV cache)
- Source: LM Studio documentation, server mode without a UI
#Frequently asked questions
Is a Mac mini M4 with 16 GB enough for a local LLM?+
What is the best LLM for a 16 GB Mac mini M4?+
Mac mini M4 or M4 Pro for local AI?+
Is the M6 Mac mini much faster than the M4 for an LLM?+
Can you connect multiple Mac minis in a cluster for an LLM?+
How much power does an M4 Mac mini use as an LLM server 24 h/24?+
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.