Advanced 11 minllama.cpp

Compile llama.cpp with Metal

Direct response

On a Apple Silicon Mac, llama.cpp compiles in three commands and Metal is enabled by default: clone the repository, run cmake -B build, then cmake --build build --config Release, without installing a GPU toolkit. You can also install it without compiling via Homebrew. The llama-cli and llama-server binaries then place computation on the M-series chip’s GPU; unified memory, more than raw power, determines the possible model size.

llama.cpp is the inference engine underlying Ollama and LM Studio, and it runs natively on Mac. This guide shows how to install or compile it with Metal, run it on a GGUF model, serve it through an API, determine macOS's GPU memory limit, and correctly read published benchmarks for M1 through M5 chips.

By Mohamed Meguedmi·Update 2026-09-30·Tested on macOS 14+

#llama.cpp on Mac: what Metal brings

On macOS, the GPU is used through Metal, Apple's graphics and compute API, and llama.cpp uses it directly. The project's README states that Apple Silicon is a first-class citizen, optimized through ARM NEON, Accelerate, and Metal. In practice, this means the default build already produces an engine that uses the GPU, with no option to add, and unified memory avoids any transfer between the CPU and GPU: the model exists in memory only once. The limiting factor on a Mac is therefore this memory's capacity and bandwidth.

→
Nothing to install on the GPU side
Unlike CUDA, which requires a toolkit several gigabytes in size, Metal is provided by macOS. You only need the Apple build tools and CMake.

#Prerequisites

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates
An Apple Silicon Mac
M1 through M5: that's the target of this guide. Intel-based Macs gain almost nothing from it.
The compilation tools Apple
Install the Command Line Tools with xcode-select --install.
CMake and Git
Available via Homebrew: brew install cmake git.
Memory for the model
The site’s rule of thumb: an 8B in Q4 weighs about 5 GB, a 14B about 9 GB, and a 32B 19 to 20 GB, before context.
Build tools
xcode-select --install
brew install cmake git

#Install without compiling: Homebrew

If you do not need special compilation options, Homebrew is the fastest route. The llama.cpp installation documentation states that the formula is automatically updated with each new project release. You get the same ready-to-use executables, with Metal.

Installation with Homebrew
brew install llama.cpp

Compile it yourself when you want the latest repository version, need to test a branch, or want to modify a compilation option. Otherwise, Homebrew is enough: you save compilation time, and updates are handled with brew upgrade.

#1. Compile with Metal

Compilation (Metal enabled by default)
git clone https://github.com/ggml-org/llama.cpp
cd llama.cpp

cmake -B build
cmake --build build --config Release -j $(sysctl -n hw.ncpu)

The build documentation is explicit: on macOS, Metal is enabled by default, and it makes computation run on the GPU. No option needs to be added. The old repository name under Georgi Gerganov's account redirects to the ggml-org organization, where the project now resides. Binaries appear in build/bin, including llama-cli for the terminal and llama-server for the API.

i
Metal and MPS, two different things
MPS (Metal Performance Shaders) is the layer PyTorch uses. llama.cpp doesn't depend on it: its Metal backend is written for quantized-model inference, which largely explains why it's faster than solutions that use PyTorch on Mac.

Two useful options: the documentation states that -DGGML_METAL=OFF disables Metal at compile time, and that a binary compiled with Metal can be forced to use the processor with --n-gpu-layers 0, which is useful for comparing speeds.

#2. Run your first model

Recent engines can download the model themselves. The -hf option takes a Hugging Face repository name, with the quantization as a suffix; without a suffix, Q4_K_M is selected by default. For an already downloaded file, use -m with its path.

Download and run a model
./build/bin/llama-cli -hf UTILISATEUR/MODELE-GGUF:Q4_K_M
Launch a local file
./build/bin/llama-cli -m ~/modeles/mon-modele-Q4_K_M.gguf

The number of layers placed on the GPU is set with -ngl; its default value is auto, which works well on a Mac with unified memory. You can also write -ngl all to load everything. Don't look for a model path in the Ollama directory: its files are stored in an internal format that is not directly usable as a GGUF.

#Serve an OpenAI-compatible API with llama-server

llama-server exposes OpenAI API-compatible routes for chat, responses, and embeddings. By default, it listens on 127.0.0.1, port 8080: it is reachable only from your Mac. To expose it to the network, you must explicitly pass --host, and you should then protect access.

Start the local API
./build/bin/llama-server -m ~/modeles/mon-modele-Q4_K_M.gguf --port 8080

#4. macOS GPU memory limit

On Apple Silicon, macOS does not give the GPU all of the unified memory. A model that exceeds this share is rejected or spills over to the CPU, even if total memory is sufficient. The engine displays the effective value at startup: look for the line ggml_metal_init: recommendedMaxWorkingSetSize in the llama-cli or llama-server logs.

The sysctl command iogpu.wired_limit_mb lets you raise this limit. It expects a value in megabytes: 61 440 for 60 GB, for example. A llama.cpp repository contributor notes that it must be reapplied at every startup because it is not persistent, and advises against going as high as 100%: the system needs memory for everything not locked by the GPU, and things go badly if you do not leave it enough.

Raise the limit (not persistent)
# 56 Go pour le GPU sur un Mac de 64 Go (56 x 1024 = 57344)
sudo sysctl iogpu.wired_limit_mb=57344

# Relancez ensuite le modèle et relisez recommendedMaxWorkingSetSize
!
Leave room for macOS
Leave the system at least a few gigabytes. A limit set too high causes slowdowns or freezes that require a restart. To make the setting permanent, you must rerun it at startup, for example with a daemon launched at boot: there is no native persistent setting.

#Which model for which unified memory

Unified memory is shared between macOS, your applications, and the model, and the GPU receives only part of it. The site’s guideline places the weights of an 8B model in Q4 at around 5 GB, a 14B model at around 9 GB, and a 32B model at around 19 to 20 GB. The context cache is added on top. The table provides a cautious rule of thumb; the authoritative value is recommendedMaxWorkingSetSize, displayed by the engine.

Approximate scale by Mac memory (Q4 sizes, excluding context)
Mac memoryReasonable modelNote
8 GB3B (2 GB)An 8B model is possible but leaves too little headroom on macOS
16 GB8B (5 GB), medium contextThe default GPU ceiling remains sufficient
24 to 32 GB14B (9 GB), or even an 8B in Q8A 32B in Q4 requires raising the GPU limit
48 to 64 GB32B (19–20 GB) with a long contextRaise the GPU limit if the engine rejects the model
96 GB and more70B (about 40 GB) and MoE modelsCheck recommendedMaxWorkingSetSize before downloading

These orders of magnitude are cautionary estimates, not measurements: a long context, a second model, or a demanding application is enough to change the picture. The dedicated memory guide provides the complete calculation method.

#5. The options that matter

llama-cli and llama-server options (official llama-cli README)
OptionRoleDefault value
-ngl, --n-gpu-layersNumber of layers placed in VRAM (a number, auto, or all)auto
-fa, --flash-attnFlash Attention: on, off, or autoauto
-ctk, -ctvKV cache type for keys and values (f16, q8_0, q4_0…)f16
-hfHugging Face repository to download, with optional quantizationQ4_K_M if the suffix is omitted
-cContext size, in tokensdepending on the model

Flash Attention is enabled automatically by default, so in most cases there is no need to enable it manually. The KV cache can be quantized with -ctk and -ctv, provided Flash Attention is active; compared with f16, q8_0 roughly halves the cache memory usage, at the cost of a slight loss of precision that you should verify for your use cases. The dedicated guide explains the tradeoff in detail.

#6. Performance by chip: what public benchmarks measure

QuelLLM does not benchmark these machines. The reference is the “Performance of llama.cpp on Apple Silicon M-series” discussion in the llama.cpp repository, where each contributor runs the same test on a 7B LLaMA in Q4_0. The table below includes a few entries, along with the llama.cpp versions used for each measurement: those for the M1 through M4 chips are from the same version, while the M5 version is newer.

Generation (tg) and prompt (pp) on LLaMA 7B Q4_0, in tokens per second
Chip (GPU cores)BandwidthPromptGenerationShare of the theoretical ceiling
M2 Pro (19)200 GB/s341,1938,8674 %
M3 Pro (18)150 GB/s341,6730,7478 %
M4 Pro (20)273 GB/s439,7850,7471 %
M5 Pro (20)307 GB/s1 620,6466,3382 %
M4 Max (40)546 GB/s885,6883,0658 %

The theoretical ceiling is bandwidth divided by model weight (3.56 GiB, or 3.82 GB). Pro chips reach 71 to 82% of that ceiling, while the Max chip reaches only 58%: beyond a certain point, memory is no longer the only bottleneck, and paying for more bandwidth yields less than the spec sheet suggests.

Three takeaways. Generation follows bandwidth: the M3 Pro, at 150 GB/s, is slower than the M2 Pro at 200 GB/s, despite being a newer chip generation. Max chips, with much higher bandwidth, dominate. Finally, prompt processing took a leap with the M5: 1,620.64 tokens/s for an M5 Pro, versus 439.78 for an M4 Pro with the same number of GPU cores, or 3.7 times faster. This gap matters for long documents and RAG, much less for chat.

i
How to read these figures
The measurements come from different contributors, using different versions of llama.cpp and macOS. They provide a rough basis for comparing chips, not a guarantee for your machine. An 8- to 9-billion-parameter model in Q4 is heavier than a 7B and will therefore run somewhat more slowly.

A good habit before concluding that a Mac is slow: run the same model with -ngl 0, then with the default value, and compare. The difference between the two shows what the GPU actually contributes on your machine and confirms that computation is going through Metal. Also note the llama.cpp version in use: Metal optimizations evolve quickly, and an old binary can be significantly slower than a recent version.

#Troubleshooting: common errors

Common symptoms on Mac
SymptomLikely causeTrack
Very low throughput, processor at 100%Model loaded on the CPUCheck -ngl and read the startup logs
GPU memory allocation errorModel larger than the share of RAM allocated to the GPUIncrease iogpu.wired_limit_mb, reduce the model or context
The Mac slows down or freezesGPU ceiling too high, no headroom for macOSLower the value of iogpu.wired_limit_mb
Compilation failsMissing Apple tools or CMake is too oldxcode-select --install then brew upgrade cmake
The model was not foundIncorrect Hugging Face repository path or nameTest -hf with a known repository, or -m with an absolute path
FAQ
Do you need to compile llama.cpp to use Metal on Mac?+
No. Metal is enabled by default when compiling on macOS, and Homebrew provides ready-to-use executables with brew install llama.cpp. Compile it yourself only for the latest repository version, a branch you want to test, or a specific option. Both approaches provide the same tools, llama-cli and llama-server.
How can I check whether llama.cpp is using the GPU on my Mac?+
Read the logs when launching llama-cli or llama-server: they show Metal initialization and the value of recommendedMaxWorkingSetSize. Very low throughput with the processor saturated indicates that a model is running on the CPU. In that case, check the -ngl setting, whose default value is auto.
How do you give a GPU more memory on a Mac Apple Silicon?+
With sudo sysctl iogpu.wired_limit_mb=VALEUR, expressed in megabytes. The setting is not persistent and must be reapplied at every startup. Do not aim for 100% of RAM: macOS needs memory for the rest of the system. Check recommendedMaxWorkingSetSize in the logs to confirm the new value.
llama.cpp or Ollama on Mac?+
Ollama is built on llama.cpp and simplifies installation, downloads, and the API. Use llama.cpp directly for the latest options, fine-grained parameter control, or llama-server. On Mac, the comparative MLX versus llama.cpp guide also covers the other possible engine. To get started, Ollama or LM Studio are enough.
What speed should you expect from a Mac M4 Pro with llama.cpp?+
The public llama.cpp benchmark reports 50.74 tokens/s for generation and 439.78 for prompt processing on a 20-GPU-core M4 Pro running a LLaMA 7B in Q4_0. A larger model will be slower. These figures vary with the llama.cpp version, memory, and macOS.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.