Beginner 14 minMac mini

Which LLM on a Mac mini M4 / M4 Pro (16–64 GB) ?

Direct response

A Mac mini M4 runs a local LLM well starting at 16 GB: models with 7 to 12 billion parameters, at best around twenty tokens/s on a 7B in Q4_0, because its 120 GB/s of bandwidth limits speed. To target 14 to 35 billion parameters, you need an M4 Pro (273 GB/s, 24 to 64 GB). Apple replaced this lineup in August 2026 with the Mac mini M6 and M5 Pro.

This guide puts the figures from Apple, llama.cpp, and the Ollama and LM Studio documentation to work on one decision: which chip, how much memory, which model, and what speed to expect. It distinguishes what Apple claims, what the community has measured, and what is merely a calculation, and reviews the M6 and M5 Pro lineup released in September 2026.

Choosing a machine? Our picks by budget →

By Mohamed Meguedmi·Update 2026-09-29·Tested on macOS 14+
Recommended hardware

Buying alternative for this guide: GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395).

A mini PC is a complete machine: check the available memory and engine compatibility. It does not replace macOS/MLX or CUDA.

Why this choice? Our complete guide on GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395) →

Compare all options by budget, from €800 to €3,500 →

On the go: which laptop for local AI →

Affiliate links — possible commission at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.

#Mac mini M4 and local LLMs: what the machine can do

A Mac mini makes a good local LLM server because its unified memory is shared between the CPU and GPU, with no separate VRAM. Two figures determine what you can do with it. Capacity (16 to 64 GB depending on the chip) determines the model size that can load: 7 to 12 billion parameters on 16 GB, and 24 to 35 billion on 32 to 48 GB. Bandwidth determines speed: 120 GB/s for the M4 chip and 273 GB/s for the M4 Pro, according to Apple. An LLM rereads its weights for every generated token, so maximum throughput is roughly the bandwidth divided by the weight size. For a 7B model in Q4_0, llama.cpp's public table reports 24.1 tokens/s for an M4 and about 50 for an M4 Pro. Apple introduced Mac mini M6 and M5 Pro models in August 2026 that replace this lineup; the M4 was still available from some retailers in late August.

Mac mini M4
10-core CPU, 10-core GPU, 120 GB/s, 16 GB unified memory (24 or 32 GB optional), Thunderbolt 4.
Mac mini M4 Pro
12-core CPU, 16-core GPU (14 and 20 optional), 273 GB/s, 24 GB of memory (48 or 64 GB optional), Thunderbolt 5.
Format and noise
12.7 × 12.7 × 5.0 cm, 0.67 kg (M4) or 0.73 kg (M4 Pro); 5 dBA at idle, according to Apple.
Advertised consumption
M4: 4 W at idle, 65 W at maximum. M4 Pro: 5 W and 140 W. Apple does not publish an inference figure.
!
The M4 lineup has been replaced since August 2026
Apple introduced the M6 and M5 Pro Mac mini in August 2026, available starting September 22. The tables below apply to an M4 already purchased, in stock, or used.

#Mac mini M4 or M4 Pro: bandwidth makes the difference

The Mac Kit

You know which model fits on your M4 Mac mini. The Mac kit teaches you how to get the most out of it: push the GPU memory limit further (Ch. 2), choose between MLX and GGUF (Ch. 3), and turn your Mac mini into an AI server for the whole house (Ch. 12).

  • Lifetime online access
  • PDF + files
  • Lifetime updates

Between the two chips, the real difference is not maximum memory (32 GB versus 64 GB) but bandwidth, multiplied by 2.3 (273 ÷ 120). In the llama.cpp table, the measured ratio is 2.1 for generation (50.7 ÷ 24.1 in Q4_0) and 2.0 for prompt processing (440 ÷ 221).

Mac mini M4 and M4 Pro for an LLM (sources: Apple, llama.cpp)
CriterionMac mini M4Mac mini M4 Pro
Memory bandwidth (Apple)120 GB/s273 GB/s
Unified memory (Apple)16 GB, with 24 or 32 GB options24 GB, with a 48 or 64 GB option
Generation, Llama 2 7B in Q4_0 (llama.cpp)24.1 tokens/s50.7 tokens/s (20-core GPU)
Generation, same model in Q8_013.5 tokens/s30.7 tokens/s
Prompt processing, Q4_0221 tokens/s440 tokens/s
Thunderbolt ports (Apple)Thunderbolt 4, 40 Gb/sThunderbolt 5, 120 Gb/s
The M4 is enough if
you are on your own, target models with up to 12 billion parameters (Qwen 3.5 9B, Gemma 4 12B) and accept around twenty tokens/s at best. Choose 24 GB to keep an embeddings model loaded as well.
The M4 Pro is the clear choice if
you are targeting models with 14 to 27 billion parameters, a 30- to 35-billion-parameter MoE (48 GB minimum), a coding agent, or multiple users. Beyond 64 GB, you need a Mac Studio.

#VRAM on Mac mini: unified memory and its GPU ceiling

A Mac mini has no dedicated VRAM, but the GPU cannot use all the memory. According to llama.cpp users, the default ceiling is two-thirds to three-quarters of the RAM; at launch, llama.cpp displays the recommendedMaxWorkingSetSize line, which gives the actual value. The iogpu.wired_limit_mb setting defaults to 0, meaning “system default policy,” not “unlimited.”

Default GPU limit, estimated at two-thirds and three-quarters of RAM
Mac mini memoryEstimated GPU ceilingRemains for macOS
16 GB10.7 to 12 GB4 to 5 GB
24 GB16 to 18 GB6 to 8 GB
32 GB21 to 24 GB8 to 11 GB
48 GB32 to 36 GB12 to 16 GB
64 GB43 to 48 GB16 to 21 GB

The ceiling must hold the weights, the KV cache (the context memory, which grows with the conversation), and compute buffers. To load a model that is close to the limit, raise the ceiling: MLX recommends a value higher than the model size in megabytes but lower than the machine’s memory. Leave several gigabytes for macOS, or speed will collapse. A manual sysctl setting may not survive a reboot.

Read it, then note the ceiling (example: M4 Pro 48 GB)
# Valeur actuelle : 0 = plafond par défaut du système
sysctl iogpu.wired_limit_mb

# 40 Go pour le GPU sur 48 Go de RAM (40 x 1024 = 40960 Mo)
sudo sysctl iogpu.wired_limit_mb=40960

#Which model for which memory: what fits in a Mac mini

The table compares the file sizes published by Ollama with a cap of two-thirds of RAM (a conservative assumption). “Yes”: the files use less than 70% of the cap, leaving room for the KV cache. “Barely”: they fit, with a short context. “No”: they exceed it. These are calculations, not measurements.

Default quantization of Ollama unless noted: does the model fit within the default GPU limit?
ModelWeights (Ollama)M4 16 GBM4 24 GBM4 32 GBM4 Pro 48 GBM4 Pro 64 GB
Qwen 3.5 9B6.6 GByesyesyesyesyes
Gemma 4 12B7.6 GBjustyesyesyesyes
gpt-oss 20B14 GBnojustyesyesyes
Mistral Small 24B14 GBnojustyesyesyes
Qwen 3.8 27B18 GBnonojustyesyes
Qwen3-Coder 30B-A3B19 GBnonojustyesyes
Qwen 3.6 35B-A3B (MoE)23 GBnononojustyes
Qwen3-Coder 30B-A3B in Q8_032 GBnononojustjust

#Mac mini M4 16 GB: models to prioritize

On 16 GB, the ceiling is around 11 GB. Qwen 3.5 9B (6.6 GB at Ollama) leaves about 4 GB for the KV cache and buffers. Gemma 4 12B (7.6 GB) works with a short context. A 14 GB model such as Mistral Small 24B overflows: Ollama splits it between the GPU and CPU, and speed drops. The ollama ps command displays the split; aim for 100% GPU.

→
An MoE reduces time per token, not memory
Qwen3-Coder 30B-A3B has 30.5 billion parameters in total but 3.3 billion active per token: it reads relatively few weights and runs quickly, but the entire file (19 GB) remains in memory. It therefore does not fit on a 16 GB Mac mini despite its 3 billion active parameters.

#Mac mini speed: theoretical ceiling and published measurements

Generation throughput is limited by bandwidth: tokens per second ≈ bandwidth (GB/s) ÷ weights read per token (GB). This calculation gives an upper bound, never a measurement. For a dense 14 GB model, 120 ÷ 14 gives 8.6 tokens/s at best on an M4, and 273 ÷ 14 gives 19.5 on an M4 Pro.

Theoretical ceiling in tokens/s (bandwidth ÷ Ollama weights, dense models)
ModelWeightsM4 (120 GB/s)M4 Pro (273 GB/s)
Qwen 3.5 9B6.6 GB1841
Gemma 4 12B7.6 GB1636
Mistral Small 24B14 GB8,619,5

The published measurements remain below this ceiling. Here are the figures from the table that the llama.cpp community maintains for the Apple chips, with Llama 2 7B. These are not our measurements, and they cover an older model in Q4_0 and Q8_0, but they show the performance ratio between the chips.

llama.cpp, Llama 2 7B: tokens/s (GitHub discussion 4167, commit 8e672ef)
ChipBandwidthQ4_0 generationQ8_0 generationQ4_0 prompt
M4 (10-core GPU)120 GB/s24,113,5221
M4 Pro (16-core GPU)273 GB/s49,630,5364
M4 Pro (20-core GPU)273 GB/s50,730,7440
M5 Pro (20-core GPU)307 GB/s66,338,91 621

Moving from Q8_0 to Q4_0 speeds up generation by a factor of 1.8 on M4 and 1.7 on M4 Pro: file size matters more than computation. GPU cores matter mainly for prompt processing (364 versus 440 tokens/s on M4 Pro). The table contains no M6 measurements as of the date consulted.

i
A reasoning model takes time to respond
A model that reasons often produces 2,000 tokens before its answer. At 8 tokens/s (a dense 14B on M4), that takes 250 seconds, or more than 4 minutes. At 19 tokens/s on M4 Pro, it takes about 105 seconds. This calculation matters when choosing between M4 and M4 Pro.

#Mac mini M6 and M5 Pro: what the new lineup changes for an LLM

Apple announces up to 4.8 times faster prompt processing on the Mac mini M6 in LM Studio than on the M4, and up to 4 times faster on the M5 Pro than on the M4 Pro. These gains apply to prompt processing, which benefits from the neural accelerators added to each GPU core. Generation remains governed by memory bandwidth and improves much less: from 12% (307 versus 273 GB/s) to 42% (170 versus 120 GB/s).

The four chips (source: Apple)
ChipBandwidthUnified memoryIdle / maximum
M4120 GB/s16, 24, or 32 GB4 W / 65 W
M6153 GB/s (16 GB), 170 GB/s (24 GB and up)16, 24, or 32 GB4 W / 70 W
M4 Pro273 GB/s24, 48, or 64 GB5 W / 140 W
M5 Pro307 GB/s24, 48, or 64 GB6 W / 145 W

A base M6 at 153 GB/s has only 28% more bandwidth than the M4: the theoretical ceiling for a 9B rises from about 18 to 23 tokens/s. The new lineup is mainly beneficial for long-context use cases: RAG that injects long documents, or a coding agent that rereads a repository.

Nine, 30- to 35-billion-parameter model
An M5 Pro with at least 48 GB, or 64 GB to work comfortably with a 35-billion-parameter MoE: the M6 tops out at 32 GB.
M4 still in stock
A reasonable choice for 7 to 12 billion parameters if the price gap with the M6 is significant. This stock is limited.

#Install an LLM server on a Mac mini with Ollama

Ollama on macOS works like an application, with its settings configured through environment variables defined with launchctl. By default, Ollama listens only on 127.0.0.1, port 11434: without changes, no other device can use it.

  1. 01
    Install the Ollama app
    Download Ollama from ollama.com and place the app in Applications. On first launch, it offers to create the ollama command link.
  2. 02
    Open the port to the local network
    Set OLLAMA_HOST with launchctl (command below), then restart the application. The address 0.0.0.0 opens the API to the entire network, without authentication: limit it to a trusted network.
  3. 03
    Prevent sleep
    In System Settings, under Energy, enable the option that prevents automatic sleep when the display is off; without it, the service may stop.
  4. 04
    Download a suitable model
    In the table, choose a model marked “yes” for your memory, such as qwen3.5:9b on 16 GB.
  5. 05
    Check from another device
    Send a request to the Mac mini’s address (.local name or IP), then run ollama ps: the Processor column should show 100% GPU.
Ollama server on the Mac mini
# Écoute sur tout le réseau local, puis relancer l'application Ollama
launchctl setenv OLLAMA_HOST "0.0.0.0:11434"

# Modèle pour 16 Go de mémoire
ollama pull qwen3.5:9b

# Test depuis un autre appareil (remplacez l'adresse)
curl http://mac-mini.local:11434/api/chat -d '{"model": "qwen3.5:9b", "messages": [{"role": "user", "content": "Bonjour"}], "stream": false}'

Two settings often copied from elsewhere are useless here: Ollama's documentation says Flash Attention is enabled automatically when the hardware allows it, and OLLAMA_ORIGINS concerns browser extensions. By contrast, quantizing the KV cache (OLLAMA_KV_CACHE_TYPE=q8_0) cuts its size roughly in half for a very small loss, which matters on 16 GB when the context grows.

#LM Studio: the alternative to Ollama

LM Studio requires an Apple Silicon chip, macOS 14 or later, and 16 GB of recommended RAM. It loads GGUF models like MLX, starting with version 0.3.4. Without a display, its llmster daemon starts with lms daemon up. The “Serve on Local Network” option exposes the API to the network; token authentication is not enabled by default. Ollama’s MLX engine, presented in preview in March 2026, required more than 32 GB: an M4 16 GB did not benefit from it.

#What a Mac mini can serve: chat, RAG, agents, cluster

Self-hosted chat server
The Ollama and LM Studio APIs accept OpenAI-format clients (Open WebUI, Zed, scripts). Placed near the router, the Mac mini serves the whole house.
Document RAG
You need to load both the generation model and an embeddings model. nomic-embed-text weighs 274 MB: on 16 GB, it fits alongside Qwen 3.5 9B.
Code agent
An agent sends long prompts at every step: prompt reading matters more than generation, and that's where M6 and M5 Pro gain the most.

#How many concurrent users?

By default, Ollama handles one request at a time per model (OLLAMA_NUM_PARALLEL is 1). Each parallel request increases the allocated context: according to the documentation, 4 requests with a context of 2,000 tokens allocate 8,000 tokens. Cache memory = cache for one conversation × parallel requests. Each user gets a slower response; no reliable measurement is available here for a specific number of users.

#Link multiple Mac minis into a cluster

Only the M4 Pro offers Thunderbolt 5 (120 Gb/s); the M4 is limited to Thunderbolt 4 (40 Gb/s). The exo project supports RDMA over Thunderbolt 5, which would reduce latency between machines by 99%, and says it speeds up models as devices are added, on macOS 26.2 or later. Its benchmarks cover Mac Studio, not Mac mini. A cluster is primarily used to load a model that is too large for one machine.

#Mac mini M4 or MacBook Air M4: same chip, different use case

The 13-inch M4 MacBook Air uses the same chip, with the same memory bandwidth (120 GB/s) and the same 16, 24, or 32 GB options: the speed ceiling is identical. What changes is the chassis. The previous version of this guide estimated a 20% slowdown for the Air after ten minutes; no source confirms this, so it has been removed. For a service running continuously, the Mac mini is the natural choice; for a mobile workstation (1.24 kg, including the display and battery), the MacBook.

#Frequently asked questions

Is a Mac mini M4 with 16 GB enough for a local LLM?+
Yes, for models with 7 to 12 billion parameters. Qwen 3.5 9B (6.6 GB) fits within the approximately 11 GB GPU limit with a reasonable context. Expect at best around twenty tokens/s: the M4 tops out at 120 GB/s. Beyond 12 billion parameters, choose 24 GB or an M4 Pro.
What is the best LLM for a 16 GB Mac mini M4?+
Qwen 3.5 9B is the safest choice: 6.6 GB from Ollama, leaving room for the KV cache. Gemma 4 12B (7.6 GB) works with a short context. Avoid models of 14 GB or more: they spill over to the CPU and lose speed.
Mac mini M4 or M4 Pro for local AI?+
The M4 Pro, if you are targeting models with more than 12 billion parameters: its 273 GB/s versus 120 GB/s roughly doubles throughput, and its 48 and 64 GB options support 30- to 35-billion-parameter MoE models. For a 9B in solo use, the M4 is sufficient.
Is the M6 Mac mini much faster than the M4 for an LLM?+
For prompt processing, yes: Apple claims up to 4.8× more speed in LM Studio. For generation, no: bandwidth increases from 120 to 153 GB/s on the base M6, or about 28% more. Short answers change little; long contexts change a lot.
Can you connect multiple Mac minis in a cluster for an LLM?+
Yes, with a tool like exo, especially for loading a model that's too large for a single machine. The M4 Pro offers Thunderbolt 5 (120 Gb/s), while the M4 has only Thunderbolt 4 (40 Gb/s). The gains published by exo are for Mac Studio systems: verify them on your configuration.
How much power does an M4 Mac mini use as an LLM server 24 h/24?+
Apple lists 4 W at idle and 65 W maximum for the M4 (5 W and 140 W for the M4 Pro). At continuous idle, 4 W × 8,760 h yields about 35 kWh per year; multiply by your price per kWh. During inference, Apple does not publish a figure: a watt meter gives you the actual value.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.