Beginner 12 minMac mini

Which LLM on Mac mini M2 / M2 Pro (8–32 GB) ?

Direct response

An M2 or M2 Pro Mac mini makes a quiet home LLM server: with 16 GB of unified memory, it runs an 8-9B model in Q4; with 32 GB (M2 Pro only), a 24 to 32B model or a 30-35B MoE. What determines usability is bandwidth—100 GB/s on the M2 versus 200 GB/s on the M2 Pro: with the same model, the M2 Pro generates almost twice as fast. Avoid the 8 GB version.

The January 2023 Mac mini M2 was replaced by the M4 in late 2024, but it remains widely available on the used market. This page explains what each memory configuration can really do, what published measurements show in tokens per second, how to install it as a Ollama server accessible across your local network, and when a newer model becomes preferable.

Choosing a machine? Our picks by budget →

By Mohamed Meguedmi·Update 2026-09-30·Tested on macOS 14+
Recommended hardware

Buying alternative for this guide: GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395).

A mini PC is a complete machine: check the available memory and engine compatibility. It does not replace macOS/MLX or CUDA.

Why this choice? Our complete guide on GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395) →

Compare all options by budget, from €800 to €3,500 →

On the go: which laptop for local AI →

Affiliate links — possible commission at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.

#M2 / M2 Pro Mac mini in 2026: what the specs promise

Apple introduced the Mac mini M2 in January 2023 with “up to 24 GB of unified memory and 100 GB/s of bandwidth,” and the M2 Pro with 200 GB/s, twice as much, and up to 32 GB of memory. The Apple technical specifications list the standard and custom configurations: the M2 starts at 8 GB, configurable to 16 or 24 GB; the M2 Pro starts at 16 GB, configurable to 32 GB. So there is neither an M2 with 32 GB nor an M2 Pro with 8 GB. The memory is soldered: you choose it at purchase, not afterward.

Mac mini M2
8 CPU cores, 10 GPU cores, 100 GB/s, and 8, 16, or 24 GB of unified memory.
Mac mini M2 Pro
10 or 12 CPU cores, 16 or 19 GPU cores, 200 GB/s, 16 or 32 GB.
Cooling
Active fan: Apple advertises a thermal system designed for sustained performance, which matters for a machine answering queries all day.
Power consumption
According to the official Apple specification sheet: 7 W at idle, 50 W maximum for the M2, and 100 W maximum for the M2 Pro. The lower values sometimes reported (10 to 30 W) are user measurements, not limits.
i
Why the M2 remains viable as a server
The main advantage is not raw speed but the balance of noise, size, and idle power consumption: 7 W at idle according to Apple. A machine that responds for ten minutes a day and sleeps the rest of the time costs very little to run. For speed, look at memory bandwidth.

#M2 or M2 Pro: the choice comes down to bandwidth

The Mac Kit

Local AI on your Mac, done right: the models that fit your unified memory, MLX vs Ollama vs LM Studio, and the settings that make Apple Silicon fast.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

During text generation, each token forces the chip to reread all of the model's active weights from memory. Speed is therefore capped by bandwidth divided by model size. An M2 at 100 GB/s and a 5 GB model yield a theoretical ceiling close to 20 tokens per second; the M2 Pro, at 200 GB/s, doubles that ceiling. Memory capacity only determines what fits: a 24B model in Q4 (about 14 GB) leaves nothing for the system on a 16 GB machine.

M2 or M2 Pro for an LLM server
CriterionMac mini M2Mac mini M2 Pro
Bandwidth100 GB/s200 GB/s
Possible memory8, 16, or 24 GB16 or 32 GB
Largest comfortable use case8–9B in Q4 at 16 GB; 12–14B at 24 GB24B to 32B in Q4 at 32 GB
Multi-user, RAGOne or two light usersSeveral short requests are acceptable
Maximum consumption (Apple)50 W100 W

The 24 GB M2 is the hybrid configuration: it has the capacity of a 14B, but retains the M2's memory bandwidth. It can therefore load larger models than the 16 GB M2 without running them faster. Choose it if you want a long context window or two loaded models; otherwise, the 16 GB M2 Pro, twice as fast with the same model, is often better for interactive use.

#Which model for which memory: the calculation, not the promise

Not everything fits in GPU memory. On Apple Silicon, Metal limits by default how much unified memory the GPU can use, according to discussions in the llama.cpp repository, to between two-thirds and three-quarters. On a 16 GB machine, allow about 10 to 12 GB for the model and its context, with the rest left to macOS. The site's reference point for Q4 weights is 8B ≈ 5 GB, 14B ≈ 9 GB, 32B ≈ 19–20 GB; the context cache is added on top.

What fits based on memory (Q4 weights from the QuelLLM catalog, excluding context)
Model (Q4)WeightsM2 8 GBM2 16 GBM2 24 GB / M2 Pro 16 GBM2 Pro 32 GB
Qwen 3.5 4B2.3 GBYesYesYesYes
Granite 4.2 8B4.6 GBJustYesYesYes
Qwen 3.5 9B6 GBNoYesYesYes
Gemma 4 12B7 GBNoJustYesYes
Mistral Small 3.2 24B14 GBNoNoNo (24 GB: very tight)Yes
Qwen 3.6 35B-A3B (MoE)21 GBNoNoNoYes, little headroom

How to read the table: “Just” means the weights fit, but the long context and macOS are competing for the remaining memory. The 8 GB option should be ruled out for any serious use: macOS and the browser already take half the memory. The MoE 35B-A3B is a special case: it weighs 21 GB but activates only about 3 billion parameters per token, making it considerably faster than a dense model of the same size. That is the best reason to choose an M2 Pro 32 GB.

#Measured throughput: what the published source allows us to say

We don't have in-house measurements and don't present any. The most useful public reference is the “Performance of llama.cpp on Apple Silicon M-series” discussion in the llama.cpp repository, where users publish text-generation results (tg128) for a Llama 7B model on each Apple chip. These figures date from the llama.cpp version available at the time and use an older model than today's models, but they provide a sense of scale.

Text generation, Llama 7B, llama.cpp (discussion #4167)
ChipBandwidthQ4_0Q8_0F16
M2, 10 GPU cores100 GB/s21,91 t/s12,21 t/s6,72 t/s
M2 Pro, 16 GPU cores200 GB/s37,87 t/s22,70 t/s12,47 t/s
M2 Pro, 19 GPU cores200 GB/s38,86 t/s23,01 t/s13,06 t/s

This table confirms three things. First, speed tracks model size: moving from Q4_0 to Q8_0 cuts throughput by approximately 1,8. Second, the M2 Pro is about 1,7 times faster than the M2, slightly less than the 2× bandwidth advantage. Third, the additional GPU cores in the 19-core M2 Pro barely affect generation (38,86 versus 37,87 t/s): they help with prompt processing, not generation.

→
The calculation to redo yourself
For an M2 at 100 GB/s and a 3.8 GB file, the theoretical ceiling is close to 26 t/s; the measurement above (21.91 t/s) represents about 84% of it. For a roughly 6 GB 9B model in Q4, the same calculation gives a ceiling close to 16 t/s on M2 and 33 t/s on M2 Pro, before losses. These are ceilings based on published measurements, not guaranteed throughput.

#Context and prompt processing: the real bottleneck

Tokens per second don’t tell the whole story. Two figures affect the experience: prompt processing time (the pp512 column from the same measurements: 201 t/s in F16 for the M2, 384 t/s for the 19-core M2 Pro) and context-cache growth. A 20,000-token input document takes several dozen seconds before the first word, on any Mac from this generation. For RAG or document analysis, keep excerpts short instead of pasting an entire PDF.

The context cache also uses memory: the longer the window, the less remains for weights. On 16 GB, reducing the window to 8,000 or 16,000 tokens leaves room for a larger model. The context-window guide explains the rule, and the KV-cache guide explains how to quantify it.

#Install Ollama as a 24/7 server: the procedure that works

Ollama requires macOS Sonoma (14) or later and listens by default on 127.0.0.1, port 11434—in other words, only from the machine itself. To make it accessible from other computers, you need to change the listening address. On Mac, the official FAQ documentation says that when Ollama runs as an application, you set the environment variables with launchctl, then restart the application.

  1. 01
    Install the Ollama app
    Download the application from the official website or use Homebrew; the macOS installation guide details both options. Do not use the application and a Homebrew service at the same time: they compete for the same port.
  2. 02
    Define the listening address
    Run launchctl setenv OLLAMA_HOST 0.0.0.0:11434 in Terminal, then quit and relaunch the Ollama application.
  3. 03
    Prevent sleep
    In System Settings, under Energy Saver, enable the option that prevents automatic sleep when the display is off, and enable automatic login if the Mac needs to restart on its own.
  4. 04
    Reserve a fixed address
    Assign a fixed IP address to the Mac mini in your router, or use its mac-mini.local name on the local network.
  5. 05
    Preload the model
    Download the model with ollama pull, then test it from another machine using the curl command below.
Configuration and testing
# Sur le Mac mini
launchctl setenv OLLAMA_HOST "0.0.0.0:11434"
# quitter puis relancer l'application Ollama
ollama pull qwen3:8b

# Depuis un autre poste du réseau local
curl http://mac-mini.local:11434/api/chat -d '{"model":"qwen3:8b","stream":false,"messages":[{"role":"user","content":"Bonjour"}]}'
!
Exposing the API: a security decision
Ollama has no authentication. Listening on 0.0.0.0 is fine for a trusted home network; for any access from outside, never forward port 11434 on your router. Use a VPN or an authenticated reverse proxy. By default, a model remains in memory for 5 minutes after the last request: adjust OLLAMA_KEEP_ALIVE if the first call of the day feels slow.

#What the Mac mini can do for the whole house

The Mac mini M2 isn’t a production server, but it’s excellent for three uses: a family chat, an OpenAI-compatible API for your tools, and lightweight automations. Ollama supports a subset of the OpenAI API on the same port, allowing you to connect many existing clients by changing only the address; verify that the features you need (tools, structured outputs) are supported.

Chat
Open WebUI or a mobile app compatible with Ollama: enter the address http://mac-mini.local:11434 in the settings.
Code
Continue, Aider, or Zed: enter the Ollama server URL as the provider.
Automation
n8n, Home Assistant, and iOS Shortcuts can call an OpenAI-compatible HTTP API.
Document indexing
A lightweight embedding model, such as nomic-embed-text, leaves room for an 8–9B chat model on 16 GB.

One limitation to know: Ollama processes one request at a time per model by default (OLLAMA_NUM_PARALLEL is 1). So two people asking a question at the same time either share the throughput or wait their turn. That's enough for a family; for a team of more than three or four active users, a more powerful machine or an inference server designed for parallelism is a better choice.

#Mac mini M2 or newer machine: when to upgrade

The base M4 increases to 120 GB/s (the value in the llama.cpp table) and the M4 Pro to 273 GB/s, according to Apple; since bandwidth is the limiting factor, the speed gap is about 20% between the base M2 and M4, and much greater between the M2 Pro and M4 Pro. A used 32 GB M2 Pro remains relevant if you are targeting a 24B to 32B model; for agentic workloads or coding with long contexts, a recent chip with more memory provides a more comfortable experience.

When to stay on M2, when to switch
SituationRecommendation
Family chat 8-9B, tight budgetA used M2 16 GB is enough
24 to 32B model or 30-35B MoEM2 Pro 32 GB, or a newer Mac with 32 GB or more
Long contexts, agents, multiple usersA Mac mini M4 Pro or newer
Requires 64 GB or moreMac Studio, never an M2 mini

#Frequently asked questions

FAQ
Le Mac mini M2 peut-il tourner 24 heures sur 24 comme serveur LLM ?+
Yes, it is designed to stay on: Apple draws 7 W at idle and has an active fan to handle sustained workloads. The key setting is to prevent automatic sleep. Under intensive inference, power consumption rises, but Apple specifies a maximum of 50 W for the M2 and 100 W for the M2 Pro.
Which used Mac mini M2 should you choose for local AI?+
The 32 GB M2 Pro is for targeting 24B to 32B models or a 30–35B MoE; the 16 GB M2 offers the best price for 8–9B models. Avoid the 8 GB model: macOS already uses half of it. Since the memory is soldered, check the exact configuration in “About This Mac” before buying.
How many tokens per second does a Mac mini M2 produce?+
Measurements published in the llama.cpp repository report 21.9 t/s on M2 and 37.9 to 38.9 t/s on M2 Pro, for a Llama 7B in Q4_0. For a 9B model in Q4, expect lower values. These are user measurements and should be treated as an order of magnitude.
Can you use the Mac mini M2 from an iPhone or another workstation?+
Yes. Once OLLAMA_HOST is set to 0.0.0.0:11434, any compatible Ollama or OpenAI client on the local network can connect to it, including iOS applications that ask for the server address. However, do not expose it to the Internet without a VPN or an authenticated reverse proxy, because the Ollama API provides no authentication.
Is the Mac mini M2 faster than a MacBook Air M2?+
With the same chip, bandwidth is identical, so peak throughput is similar. The Mac mini's advantage is its active fan: it sustains the load without reducing its clock speed, whereas a fanless MacBook Air eventually slows down on long generations.
Can you fine-tune a model on a Mac mini M2 Pro?+
For a small model, LoRA fine-tuning via MLX is possible but slow: expect several hours depending on dataset size. Full fine-tuning is not practical with 32 GB of shared memory with macOS. The most sensible approach is to use the Mac mini for inference and rent a GPU occasionally for training.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.