llama.cpp: what is it, and should you leave Ollama ?
llama.cpp is the C/C++ inference engine that runs GGUF-format models on CPU and on NVIDIA, AMD, Intel, and Apple Silicon GPUs; Ollama, LM Studio, and KoboldCpp all rely on it. Using it directly through llama-server does not make a model faster on the same hardware: it gives you control over settings that wrappers choose for you, such as GPU/CPU allocation, context size, and KV-cache quantization.
llama.cpp est le moteur d'inférence écrit en C et C++ qui exécute les modèles au format GGUF sur CPU, sur GPU NVIDIA, AMD et Intel, et sur Mac. Au 20 septembre 2026, la plupart des applications de LLM local s'appuient sur lui ou sur sa bibliothèque ggml : Ollama, LM Studio, KoboldCpp, Jan. L'utiliser directement ne rend pas votre modèle plus rapide. Cela vous donne la main sur des réglages que les surcouches décident à votre place.
#llama.cpp in brief
The project was launched in March 2023 by Georgi Gerganov, with a simple goal: run Meta’s LLaMA model on a MacBook without Python or heavy dependencies. It is released under the MIT license. Three years later, the repository (now under the ggml-org organization) supports hundreds of architectures, and the file format it defined in August 2023, GGUF, has become the de facto standard for distributing a quantized model.
Two technical choices explain this success. The first is quantization: llama.cpp can run weights compressed to 2 to 8 bits, allowing an 8-billion-parameter model to fit in 5 GB instead of 16. The second is splitting work between the CPU and GPU: when a model does not fit entirely in VRAM, some layers remain in RAM and the rest go to the GPU. It is slower than loading everything at once, but it works, and few engines offer it.
#What's in the box
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
- llama-cli
- The command-line chat. Useful for testing a model or setting in a few seconds.
- llama-server
- An HTTP server with an OpenAI-compatible API and an integrated web interface, on port 8080. This is the component that most advanced users keep running continuously.
- llama-bench
- The measurement tool. It reports prompt-reading speed and generation speed separately, allowing you to compare two configurations without fooling yourself.
- llama-quantize
- Converts a full-precision GGUF into a lighter quantization (Q4_K_M, Q5_K_M, Q8_0).
- llama-server: setting up a local OpenAI API, step by step
- GGUF and safetensors: understanding model formats
- Official documentation for the GGUF format — Hugging Face
#What the overlays add—and hide
Ollama, LM Studio and KoboldCpp provide what llama.cpp does not: a model catalog, one-command downloads, an interface, and automatic loading and unloading. In exchange, they set default values. The table below does not compare performance; it shows who decides what.
| Setting | llama.cpp nu | Ollama | LM Studio | KoboldCpp |
|---|---|---|---|---|
| Context length | -c option, free | Conservative default value, configurable through a variable or Modelfile | Per-model slider | Launch option |
| Quantization selection | Any GGUF file | Catalog tag, Q4 by default | Proposed download list | Any GGUF file |
| Layers offloaded to the GPU | -ngl option, layer by layer | Automatic | Slider | Launch option |
| KV cache quantization | -ctk and -ctv options | Global environment variable | Advanced setting | Launch option |
| API | llama-server, OpenAI-compatible | Clean API + OpenAI compatibility | OpenAI-compatible server | Clean API + OpenAI compatibility |
| Engine updates | The same day, with every commit | With a delay | With a delay | With a delay |
That last line matters more than it may seem. When a new model architecture is released, it is first supported in llama.cpp, then in higher-level wrappers a few days or weeks later. If you want to test a model during the week it is published, this is often the only path.
#The four settings that change everything
- -ngl (n-gpu-layers)
- The number of model layers placed in VRAM. A value of 99 means “everything that fits.” If the model exceeds your VRAM, lower this number until it loads: each layer left on the CPU slows generation, but a model running at 8 tokens per second is better than one that won't start.
- -c (ctx-size)
- The allocated context window. The KV cache grows with it: on an 8B model, going from 8,000 to 32,000 tokens costs several GB of VRAM. Allocate what you need, not the maximum advertised by the model.
- -fa (flash-attn)
- Enable Flash Attention, which reduces memory use and speeds up processing of long prompts. It is also a prerequisite for quantizing the KV cache.
- --n-cpu-moe
- For MoE models, keep the experts from a certain number of layers in RAM and leave the rest on the GPU. This is what makes it possible to run a 30- or 120-billion-parameter MoE on a 16 GB card at usable speed.
- Flash Attention: enable it in llama.cpp, Ollama, and vLLM
- Quantize the KV cache to save VRAM
- Split a model across multiple GPUs with tensor-split
- Official llama-server documentation (complete list of options)
#What changed: --fit sets -ngl for you
For a long time, manually setting -ngl was the first step in every llama.cpp guide: set it to 99 to send everything to the GPU, then lower it by trial and error if loading failed. That step is no longer mandatory. The project added a --fit option, enabled by default, that automatically adjusts unspecified parameters (including -ngl) so the model fits in the available memory. On a modest card, this avoids the usual back-and-forth between a failed launch and manually lowering -ngl.
#Which backend for your card
llama.cpp is compiled, or downloaded, for a given compute backend. The right choice depends solely on your hardware. The table covers the families in our database of 90 configurations.
| Your hardware | Recommended backend | Note |
|---|---|---|
| NVIDIA GTX 10 to RTX 50, desktop and laptop | CUDA | The fastest and best-tested path. |
| AMD Radeon RX 7000 and RX 9000 | ROCm (HIP) or Vulkan | ROCm is faster when it works. Vulkan installs without friction, including on Windows. |
| AMD Radeon RX 6000 and older | Vulkan | Partial or no ROCm support depending on the GPU. |
| Apple M1 to M5 | Metal | Enabled by default in macOS binaries. All unified memory is usable. |
| Integrated Intel or AMD GPU, Intel Arc | Vulkan or SYCL | Real-world gain over CPU-only operation on small models. |
| No GPU | CPU (AVX2, AVX-512, NEON) | Works everywhere. Target models with at most 8 billion parameters. |
- Compile llama.cpp with CUDA
- Compile llama.cpp with Metal on Mac
- llama.cpp with Vulkan, the universal backend
#Our measurements with llama.cpp
llama.cpp is our benchmark’s reference engine precisely because it runs identically on every platform. Here are the generation speeds measured on Llama 3.1 8B in Q4, with a 2,048-token context, a single request, and Flash Attention enabled.
| Machine | Memory | Llama 3.1 8B Q4 |
|---|---|---|
| RTX 5090 | 32 GB | 172 tok/s |
| RTX 4090 | 24 GB | 128 tok/s |
| RTX 4070 | 12 GB | 76 tok/s |
| Mac M3 Max | 64 GB unified | 64 tok/s |
| RTX 3060 | 12 GB | 44 tok/s |
| Ryzen 7 7700, CPU only | System RAM | 7.8 tok/s |
Two takeaways. First, the gap between a RTX 3060 and a standalone processor is a factor of five to six: even an entry-level card changes the experience. Second, these figures would be practically identical in Ollama or LM Studio on the same machines, since the compute engine is the same.
#Stay with Ollama or switch to llama.cpp
| Stay on Ollama or LM Studio if… | Switch to llama.cpp if… |
|---|---|
| You want to chat with a model without reading documentation. | Your model exceeds VRAM capacity and you want to fine-tune the GPU/CPU split. |
| You frequently switch models and appreciate the built-in catalog. | You want to test an architecture released this week. |
| Your tools (Open WebUI, editor extensions) expect the Ollama API. | You’re building a durable server and want to control every option, from context to KV cache. |
#Get started in ten minutes
No need to compile anything to try it. Prebuilt binaries are published for Windows, macOS, and Linux with every release, and package managers take care of the rest. The command below downloads a small model from Hugging Face and opens the web interface at http://localhost:8080.
For a more capable model from the catalog, download the GGUF file of your choice and point to it with the -m option. Compiling from source is only useful for enabling a specific backend or tracking day-to-day development.
#Most common loading errors
Most issues with llama.cpp boil down to three causes: running out of memory, a GGUF file incompatible with the installed build, or a misspelled option. The command-line error message almost always indicates which one, provided you read it to the end instead of closing it and rerunning the command.
| Message or symptom | Likely cause | Try it |
|---|---|---|
| cudaMalloc failed: out of memory | The model, with the requested context, exceeds the available VRAM | Reduce -c, let --fit adjust -ngl, or choose a lighter quantization |
| unknown model architecture | The GGUF file uses a newer architecture than the installed binary | Update to the latest published version; new architectures arrive in llama.cpp first |
| Very slow generation despite a recent GPU | Not all layers are offloaded to the GPU, often because there is not enough free VRAM. | Check the load output (the “offloaded” line) and close other applications using the card |
| error: invalid argument | An option renamed or misspelled in the command | Compare with llama-server --help, which lists the current short and long aliases |
The loading line displayed when the server starts is worth reading in full once: it shows the number of layers actually offloaded to the GPU, the effective context size, and the type of KV cache used. That's often faster than randomly adding options until it works. Also keep an eye on the installed version: llama.cpp releases nightly builds very frequently, and a fix for your card or model may already have been released without yet making it into the Homebrew or winget package you installed the previous week.
#FAQ
Is llama.cpp faster than Ollama?+
Do you need to know how to compile to use llama.cpp?+
Can you use llama.cpp without a graphics card?+
What’s the difference between llama.cpp and vLLM?+
And on Mac, llama.cpp or MLX?+
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.