Which LLM on Mac mini M2 / M2 Pro (8–32 GB) ?
An M2 or M2 Pro Mac mini makes a quiet home LLM server: with 16 GB of unified memory, it runs an 8-9B model in Q4; with 32 GB (M2 Pro only), a 24 to 32B model or a 30-35B MoE. What determines usability is bandwidth—100 GB/s on the M2 versus 200 GB/s on the M2 Pro: with the same model, the M2 Pro generates almost twice as fast. Avoid the 8 GB version.
The January 2023 Mac mini M2 was replaced by the M4 in late 2024, but it remains widely available on the used market. This page explains what each memory configuration can really do, what published measurements show in tokens per second, how to install it as a Ollama server accessible across your local network, and when a newer model becomes preferable.
Choosing a machine? Our picks by budget →
Buying alternative for this guide: GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395).
A mini PC is a complete machine: check the available memory and engine compatibility. It does not replace macOS/MLX or CUDA.
Why this choice? Our complete guide on GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395) →
Compare all options by budget, from €800 to €3,500 →
On the go: which laptop for local AI →
Affiliate links — possible commission at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.
#M2 / M2 Pro Mac mini in 2026: what the specs promise
Apple introduced the Mac mini M2 in January 2023 with “up to 24 GB of unified memory and 100 GB/s of bandwidth,” and the M2 Pro with 200 GB/s, twice as much, and up to 32 GB of memory. The Apple technical specifications list the standard and custom configurations: the M2 starts at 8 GB, configurable to 16 or 24 GB; the M2 Pro starts at 16 GB, configurable to 32 GB. So there is neither an M2 with 32 GB nor an M2 Pro with 8 GB. The memory is soldered: you choose it at purchase, not afterward.
- Mac mini M2
- 8 CPU cores, 10 GPU cores, 100 GB/s, and 8, 16, or 24 GB of unified memory.
- Mac mini M2 Pro
- 10 or 12 CPU cores, 16 or 19 GPU cores, 200 GB/s, 16 or 32 GB.
- Cooling
- Active fan: Apple advertises a thermal system designed for sustained performance, which matters for a machine answering queries all day.
- Power consumption
- According to the official Apple specification sheet: 7 W at idle, 50 W maximum for the M2, and 100 W maximum for the M2 Pro. The lower values sometimes reported (10 to 30 W) are user measurements, not limits.
#M2 or M2 Pro: the choice comes down to bandwidth
Local AI on your Mac, done right: the models that fit your unified memory, MLX vs Ollama vs LM Studio, and the settings that make Apple Silicon fast.
- Lifetime online access
- PDF + files
- Lifetime updates
During text generation, each token forces the chip to reread all of the model's active weights from memory. Speed is therefore capped by bandwidth divided by model size. An M2 at 100 GB/s and a 5 GB model yield a theoretical ceiling close to 20 tokens per second; the M2 Pro, at 200 GB/s, doubles that ceiling. Memory capacity only determines what fits: a 24B model in Q4 (about 14 GB) leaves nothing for the system on a 16 GB machine.
| Criterion | Mac mini M2 | Mac mini M2 Pro |
|---|---|---|
| Bandwidth | 100 GB/s | 200 GB/s |
| Possible memory | 8, 16, or 24 GB | 16 or 32 GB |
| Largest comfortable use case | 8–9B in Q4 at 16 GB; 12–14B at 24 GB | 24B to 32B in Q4 at 32 GB |
| Multi-user, RAG | One or two light users | Several short requests are acceptable |
| Maximum consumption (Apple) | 50 W | 100 W |
The 24 GB M2 is the hybrid configuration: it has the capacity of a 14B, but retains the M2's memory bandwidth. It can therefore load larger models than the 16 GB M2 without running them faster. Choose it if you want a long context window or two loaded models; otherwise, the 16 GB M2 Pro, twice as fast with the same model, is often better for interactive use.
#Which model for which memory: the calculation, not the promise
Not everything fits in GPU memory. On Apple Silicon, Metal limits by default how much unified memory the GPU can use, according to discussions in the llama.cpp repository, to between two-thirds and three-quarters. On a 16 GB machine, allow about 10 to 12 GB for the model and its context, with the rest left to macOS. The site's reference point for Q4 weights is 8B ≈ 5 GB, 14B ≈ 9 GB, 32B ≈ 19–20 GB; the context cache is added on top.
| Model (Q4) | Weights | M2 8 GB | M2 16 GB | M2 24 GB / M2 Pro 16 GB | M2 Pro 32 GB |
|---|---|---|---|---|---|
| Qwen 3.5 4B | 2.3 GB | Yes | Yes | Yes | Yes |
| Granite 4.2 8B | 4.6 GB | Just | Yes | Yes | Yes |
| Qwen 3.5 9B | 6 GB | No | Yes | Yes | Yes |
| Gemma 4 12B | 7 GB | No | Just | Yes | Yes |
| Mistral Small 3.2 24B | 14 GB | No | No | No (24 GB: very tight) | Yes |
| Qwen 3.6 35B-A3B (MoE) | 21 GB | No | No | No | Yes, little headroom |
How to read the table: “Just” means the weights fit, but the long context and macOS are competing for the remaining memory. The 8 GB option should be ruled out for any serious use: macOS and the browser already take half the memory. The MoE 35B-A3B is a special case: it weighs 21 GB but activates only about 3 billion parameters per token, making it considerably faster than a dense model of the same size. That is the best reason to choose an M2 Pro 32 GB.
#Measured throughput: what the published source allows us to say
We don't have in-house measurements and don't present any. The most useful public reference is the “Performance of llama.cpp on Apple Silicon M-series” discussion in the llama.cpp repository, where users publish text-generation results (tg128) for a Llama 7B model on each Apple chip. These figures date from the llama.cpp version available at the time and use an older model than today's models, but they provide a sense of scale.
| Chip | Bandwidth | Q4_0 | Q8_0 | F16 |
|---|---|---|---|---|
| M2, 10 GPU cores | 100 GB/s | 21,91 t/s | 12,21 t/s | 6,72 t/s |
| M2 Pro, 16 GPU cores | 200 GB/s | 37,87 t/s | 22,70 t/s | 12,47 t/s |
| M2 Pro, 19 GPU cores | 200 GB/s | 38,86 t/s | 23,01 t/s | 13,06 t/s |
This table confirms three things. First, speed tracks model size: moving from Q4_0 to Q8_0 cuts throughput by approximately 1,8. Second, the M2 Pro is about 1,7 times faster than the M2, slightly less than the 2× bandwidth advantage. Third, the additional GPU cores in the 19-core M2 Pro barely affect generation (38,86 versus 37,87 t/s): they help with prompt processing, not generation.
#Context and prompt processing: the real bottleneck
Tokens per second don’t tell the whole story. Two figures affect the experience: prompt processing time (the pp512 column from the same measurements: 201 t/s in F16 for the M2, 384 t/s for the 19-core M2 Pro) and context-cache growth. A 20,000-token input document takes several dozen seconds before the first word, on any Mac from this generation. For RAG or document analysis, keep excerpts short instead of pasting an entire PDF.
The context cache also uses memory: the longer the window, the less remains for weights. On 16 GB, reducing the window to 8,000 or 16,000 tokens leaves room for a larger model. The context-window guide explains the rule, and the KV-cache guide explains how to quantify it.
#Install Ollama as a 24/7 server: the procedure that works
Ollama requires macOS Sonoma (14) or later and listens by default on 127.0.0.1, port 11434—in other words, only from the machine itself. To make it accessible from other computers, you need to change the listening address. On Mac, the official FAQ documentation says that when Ollama runs as an application, you set the environment variables with launchctl, then restart the application.
- 01Install the Ollama appDownload the application from the official website or use Homebrew; the macOS installation guide details both options. Do not use the application and a Homebrew service at the same time: they compete for the same port.
- 02Define the listening addressRun launchctl setenv OLLAMA_HOST 0.0.0.0:11434 in Terminal, then quit and relaunch the Ollama application.
- 03Prevent sleepIn System Settings, under Energy Saver, enable the option that prevents automatic sleep when the display is off, and enable automatic login if the Mac needs to restart on its own.
- 04Reserve a fixed addressAssign a fixed IP address to the Mac mini in your router, or use its mac-mini.local name on the local network.
- 05Preload the modelDownload the model with ollama pull, then test it from another machine using the curl command below.
#What the Mac mini can do for the whole house
The Mac mini M2 isn’t a production server, but it’s excellent for three uses: a family chat, an OpenAI-compatible API for your tools, and lightweight automations. Ollama supports a subset of the OpenAI API on the same port, allowing you to connect many existing clients by changing only the address; verify that the features you need (tools, structured outputs) are supported.
- Chat
- Open WebUI or a mobile app compatible with Ollama: enter the address http://mac-mini.local:11434 in the settings.
- Code
- Continue, Aider, or Zed: enter the Ollama server URL as the provider.
- Automation
- n8n, Home Assistant, and iOS Shortcuts can call an OpenAI-compatible HTTP API.
- Document indexing
- A lightweight embedding model, such as nomic-embed-text, leaves room for an 8–9B chat model on 16 GB.
One limitation to know: Ollama processes one request at a time per model by default (OLLAMA_NUM_PARALLEL is 1). So two people asking a question at the same time either share the throughput or wait their turn. That's enough for a family; for a team of more than three or four active users, a more powerful machine or an inference server designed for parallelism is a better choice.
#Mac mini M2 or newer machine: when to upgrade
The base M4 increases to 120 GB/s (the value in the llama.cpp table) and the M4 Pro to 273 GB/s, according to Apple; since bandwidth is the limiting factor, the speed gap is about 20% between the base M2 and M4, and much greater between the M2 Pro and M4 Pro. A used 32 GB M2 Pro remains relevant if you are targeting a 24B to 32B model; for agentic workloads or coding with long contexts, a recent chip with more memory provides a more comfortable experience.
| Situation | Recommendation |
|---|---|
| Family chat 8-9B, tight budget | A used M2 16 GB is enough |
| 24 to 32B model or 30-35B MoE | M2 Pro 32 GB, or a newer Mac with 32 GB or more |
| Long contexts, agents, multiple users | A Mac mini M4 Pro or newer |
| Requires 64 GB or more | Mac Studio, never an M2 mini |
- Mac mini M4 / M4 Pro: the successor
- Hardware profile: Mac mini M4 Pro
- macOS settings for Apple Silicon
- Install Ollama on macOS
- Source: Apple, Mac mini technical specifications (2023)
- Source: Apple, Mac mini power consumption and thermal dissipation
- Source: llama.cpp measurements on Apple chips
- Source: official Ollama FAQ
#Frequently asked questions
Le Mac mini M2 peut-il tourner 24 heures sur 24 comme serveur LLM ?+
Which used Mac mini M2 should you choose for local AI?+
How many tokens per second does a Mac mini M2 produce?+
Can you use the Mac mini M2 from an iPhone or another workstation?+
Is the Mac mini M2 faster than a MacBook Air M2?+
Can you fine-tune a model on a Mac mini M2 Pro?+
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.