Compile llama.cpp with Metal
On a Apple Silicon Mac, llama.cpp compiles in three commands and Metal is enabled by default: clone the repository, run cmake -B build, then cmake --build build --config Release, without installing a GPU toolkit. You can also install it without compiling via Homebrew. The llama-cli and llama-server binaries then place computation on the M-series chip’s GPU; unified memory, more than raw power, determines the possible model size.
llama.cpp is the inference engine underlying Ollama and LM Studio, and it runs natively on Mac. This guide shows how to install or compile it with Metal, run it on a GGUF model, serve it through an API, determine macOS's GPU memory limit, and correctly read published benchmarks for M1 through M5 chips.
#llama.cpp on Mac: what Metal brings
On macOS, the GPU is used through Metal, Apple's graphics and compute API, and llama.cpp uses it directly. The project's README states that Apple Silicon is a first-class citizen, optimized through ARM NEON, Accelerate, and Metal. In practice, this means the default build already produces an engine that uses the GPU, with no option to add, and unified memory avoids any transfer between the CPU and GPU: the model exists in memory only once. The limiting factor on a Mac is therefore this memory's capacity and bandwidth.
#Prerequisites
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
- An Apple Silicon Mac
- M1 through M5: that's the target of this guide. Intel-based Macs gain almost nothing from it.
- The compilation tools Apple
- Install the Command Line Tools with xcode-select --install.
- CMake and Git
- Available via Homebrew: brew install cmake git.
- Memory for the model
- The site’s rule of thumb: an 8B in Q4 weighs about 5 GB, a 14B about 9 GB, and a 32B 19 to 20 GB, before context.
#Install without compiling: Homebrew
If you do not need special compilation options, Homebrew is the fastest route. The llama.cpp installation documentation states that the formula is automatically updated with each new project release. You get the same ready-to-use executables, with Metal.
Compile it yourself when you want the latest repository version, need to test a branch, or want to modify a compilation option. Otherwise, Homebrew is enough: you save compilation time, and updates are handled with brew upgrade.
#1. Compile with Metal
The build documentation is explicit: on macOS, Metal is enabled by default, and it makes computation run on the GPU. No option needs to be added. The old repository name under Georgi Gerganov's account redirects to the ggml-org organization, where the project now resides. Binaries appear in build/bin, including llama-cli for the terminal and llama-server for the API.
Two useful options: the documentation states that -DGGML_METAL=OFF disables Metal at compile time, and that a binary compiled with Metal can be forced to use the processor with --n-gpu-layers 0, which is useful for comparing speeds.
#2. Run your first model
Recent engines can download the model themselves. The -hf option takes a Hugging Face repository name, with the quantization as a suffix; without a suffix, Q4_K_M is selected by default. For an already downloaded file, use -m with its path.
The number of layers placed on the GPU is set with -ngl; its default value is auto, which works well on a Mac with unified memory. You can also write -ngl all to load everything. Don't look for a model path in the Ollama directory: its files are stored in an internal format that is not directly usable as a GGUF.
#Serve an OpenAI-compatible API with llama-server
llama-server exposes OpenAI API-compatible routes for chat, responses, and embeddings. By default, it listens on 127.0.0.1, port 8080: it is reachable only from your Mac. To expose it to the network, you must explicitly pass --host, and you should then protect access.
#4. macOS GPU memory limit
On Apple Silicon, macOS does not give the GPU all of the unified memory. A model that exceeds this share is rejected or spills over to the CPU, even if total memory is sufficient. The engine displays the effective value at startup: look for the line ggml_metal_init: recommendedMaxWorkingSetSize in the llama-cli or llama-server logs.
The sysctl command iogpu.wired_limit_mb lets you raise this limit. It expects a value in megabytes: 61 440 for 60 GB, for example. A llama.cpp repository contributor notes that it must be reapplied at every startup because it is not persistent, and advises against going as high as 100%: the system needs memory for everything not locked by the GPU, and things go badly if you do not leave it enough.
#Which model for which unified memory
Unified memory is shared between macOS, your applications, and the model, and the GPU receives only part of it. The site’s guideline places the weights of an 8B model in Q4 at around 5 GB, a 14B model at around 9 GB, and a 32B model at around 19 to 20 GB. The context cache is added on top. The table provides a cautious rule of thumb; the authoritative value is recommendedMaxWorkingSetSize, displayed by the engine.
| Mac memory | Reasonable model | Note |
|---|---|---|
| 8 GB | 3B (2 GB) | An 8B model is possible but leaves too little headroom on macOS |
| 16 GB | 8B (5 GB), medium context | The default GPU ceiling remains sufficient |
| 24 to 32 GB | 14B (9 GB), or even an 8B in Q8 | A 32B in Q4 requires raising the GPU limit |
| 48 to 64 GB | 32B (19–20 GB) with a long context | Raise the GPU limit if the engine rejects the model |
| 96 GB and more | 70B (about 40 GB) and MoE models | Check recommendedMaxWorkingSetSize before downloading |
These orders of magnitude are cautionary estimates, not measurements: a long context, a second model, or a demanding application is enough to change the picture. The dedicated memory guide provides the complete calculation method.
#5. The options that matter
| Option | Role | Default value |
|---|---|---|
| -ngl, --n-gpu-layers | Number of layers placed in VRAM (a number, auto, or all) | auto |
| -fa, --flash-attn | Flash Attention: on, off, or auto | auto |
| -ctk, -ctv | KV cache type for keys and values (f16, q8_0, q4_0…) | f16 |
| -hf | Hugging Face repository to download, with optional quantization | Q4_K_M if the suffix is omitted |
| -c | Context size, in tokens | depending on the model |
Flash Attention is enabled automatically by default, so in most cases there is no need to enable it manually. The KV cache can be quantized with -ctk and -ctv, provided Flash Attention is active; compared with f16, q8_0 roughly halves the cache memory usage, at the cost of a slight loss of precision that you should verify for your use cases. The dedicated guide explains the tradeoff in detail.
#6. Performance by chip: what public benchmarks measure
QuelLLM does not benchmark these machines. The reference is the “Performance of llama.cpp on Apple Silicon M-series” discussion in the llama.cpp repository, where each contributor runs the same test on a 7B LLaMA in Q4_0. The table below includes a few entries, along with the llama.cpp versions used for each measurement: those for the M1 through M4 chips are from the same version, while the M5 version is newer.
| Chip (GPU cores) | Bandwidth | Prompt | Generation | Share of the theoretical ceiling |
|---|---|---|---|---|
| M2 Pro (19) | 200 GB/s | 341,19 | 38,86 | 74 % |
| M3 Pro (18) | 150 GB/s | 341,67 | 30,74 | 78 % |
| M4 Pro (20) | 273 GB/s | 439,78 | 50,74 | 71 % |
| M5 Pro (20) | 307 GB/s | 1 620,64 | 66,33 | 82 % |
| M4 Max (40) | 546 GB/s | 885,68 | 83,06 | 58 % |
The theoretical ceiling is bandwidth divided by model weight (3.56 GiB, or 3.82 GB). Pro chips reach 71 to 82% of that ceiling, while the Max chip reaches only 58%: beyond a certain point, memory is no longer the only bottleneck, and paying for more bandwidth yields less than the spec sheet suggests.
Three takeaways. Generation follows bandwidth: the M3 Pro, at 150 GB/s, is slower than the M2 Pro at 200 GB/s, despite being a newer chip generation. Max chips, with much higher bandwidth, dominate. Finally, prompt processing took a leap with the M5: 1,620.64 tokens/s for an M5 Pro, versus 439.78 for an M4 Pro with the same number of GPU cores, or 3.7 times faster. This gap matters for long documents and RAG, much less for chat.
A good habit before concluding that a Mac is slow: run the same model with -ngl 0, then with the default value, and compare. The difference between the two shows what the GPU actually contributes on your machine and confirms that computation is going through Metal. Also note the llama.cpp version in use: Metal optimizations evolve quickly, and an old binary can be significantly slower than a recent version.
#Troubleshooting: common errors
| Symptom | Likely cause | Track |
|---|---|---|
| Very low throughput, processor at 100% | Model loaded on the CPU | Check -ngl and read the startup logs |
| GPU memory allocation error | Model larger than the share of RAM allocated to the GPU | Increase iogpu.wired_limit_mb, reduce the model or context |
| The Mac slows down or freezes | GPU ceiling too high, no headroom for macOS | Lower the value of iogpu.wired_limit_mb |
| Compilation fails | Missing Apple tools or CMake is too old | xcode-select --install then brew upgrade cmake |
| The model was not found | Incorrect Hugging Face repository path or name | Test -hf with a known repository, or -m with an absolute path |
- MLX vs. llama.cpp on Mac: who wins in 2026?
- Compile llama.cpp with CUDA
- Which models for 32 GB of memory
- Source: official llama.cpp repository
- Source: llama.cpp build documentation
- Source: public benchmark on Apple Silicon
- Source: discussion about the GPU memory limit of Macs
Do you need to compile llama.cpp to use Metal on Mac?+
How can I check whether llama.cpp is using the GPU on my Mac?+
How do you give a GPU more memory on a Mac Apple Silicon?+
llama.cpp or Ollama on Mac?+
What speed should you expect from a Mac M4 Pro with llama.cpp?+
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.