Intermediate 11 minllama.cpp

llama.cpp: what is it, and should you leave Ollama ?

Direct response

llama.cpp is the C/C++ inference engine that runs GGUF-format models on CPU and on NVIDIA, AMD, Intel, and Apple Silicon GPUs; Ollama, LM Studio, and KoboldCpp all rely on it. Using it directly through llama-server does not make a model faster on the same hardware: it gives you control over settings that wrappers choose for you, such as GPU/CPU allocation, context size, and KV-cache quantization.

llama.cpp est le moteur d'inférence écrit en C et C++ qui exécute les modèles au format GGUF sur CPU, sur GPU NVIDIA, AMD et Intel, et sur Mac. Au 20 septembre 2026, la plupart des applications de LLM local s'appuient sur lui ou sur sa bibliothèque ggml : Ollama, LM Studio, KoboldCpp, Jan. L'utiliser directement ne rend pas votre modèle plus rapide. Cela vous donne la main sur des réglages que les surcouches décident à votre place.

By Mohamed Meguedmi·Update 2026-09-28·Tested on macOS 14+

#llama.cpp in brief

The project was launched in March 2023 by Georgi Gerganov, with a simple goal: run Meta’s LLaMA model on a MacBook without Python or heavy dependencies. It is released under the MIT license. Three years later, the repository (now under the ggml-org organization) supports hundreds of architectures, and the file format it defined in August 2023, GGUF, has become the de facto standard for distributing a quantized model.

Two technical choices explain this success. The first is quantization: llama.cpp can run weights compressed to 2 to 8 bits, allowing an 8-billion-parameter model to fit in 5 GB instead of 16. The second is splitting work between the CPU and GPU: when a model does not fit entirely in VRAM, some layers remain in RAM and the rest go to the GPU. It is slower than loading everything at once, but it works, and few engines offer it.

#What's in the box

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates
llama-cli
The command-line chat. Useful for testing a model or setting in a few seconds.
llama-server
An HTTP server with an OpenAI-compatible API and an integrated web interface, on port 8080. This is the component that most advanced users keep running continuously.
llama-bench
The measurement tool. It reports prompt-reading speed and generation speed separately, allowing you to compare two configurations without fooling yourself.
llama-quantize
Converts a full-precision GGUF into a lighter quantization (Q4_K_M, Q5_K_M, Q8_0).

#What the overlays add—and hide

Ollama, LM Studio and KoboldCpp provide what llama.cpp does not: a model catalog, one-command downloads, an interface, and automatic loading and unloading. In exchange, they set default values. The table below does not compare performance; it shows who decides what.

Control comparison, not speed · status as of 20/09/2026
Settingllama.cpp nuOllamaLM StudioKoboldCpp
Context length-c option, freeConservative default value, configurable through a variable or ModelfilePer-model sliderLaunch option
Quantization selectionAny GGUF fileCatalog tag, Q4 by defaultProposed download listAny GGUF file
Layers offloaded to the GPU-ngl option, layer by layerAutomaticSliderLaunch option
KV cache quantization-ctk and -ctv optionsGlobal environment variableAdvanced settingLaunch option
APIllama-server, OpenAI-compatibleClean API + OpenAI compatibilityOpenAI-compatible serverClean API + OpenAI compatibility
Engine updatesThe same day, with every commitWith a delayWith a delayWith a delay

That last line matters more than it may seem. When a new model architecture is released, it is first supported in llama.cpp, then in higher-level wrappers a few days or weeks later. If you want to test a model during the week it is published, this is often the only path.

#The four settings that change everything

-ngl (n-gpu-layers)
The number of model layers placed in VRAM. A value of 99 means “everything that fits.” If the model exceeds your VRAM, lower this number until it loads: each layer left on the CPU slows generation, but a model running at 8 tokens per second is better than one that won't start.
-c (ctx-size)
The allocated context window. The KV cache grows with it: on an 8B model, going from 8,000 to 32,000 tokens costs several GB of VRAM. Allocate what you need, not the maximum advertised by the model.
-fa (flash-attn)
Enable Flash Attention, which reduces memory use and speeds up processing of long prompts. It is also a prerequisite for quantizing the KV cache.
--n-cpu-moe
For MoE models, keep the experts from a certain number of layers in RAM and leave the rest on the GPU. This is what makes it possible to run a 30- or 120-billion-parameter MoE on a 16 GB card at usable speed.

#What changed: --fit sets -ngl for you

For a long time, manually setting -ngl was the first step in every llama.cpp guide: set it to 99 to send everything to the GPU, then lower it by trial and error if loading failed. That step is no longer mandatory. The project added a --fit option, enabled by default, that automatically adjusts unspecified parameters (including -ngl) so the model fits in the available memory. On a modest card, this avoids the usual back-and-forth between a failed launch and manually lowering -ngl.

→
-ngl 99 remains valid
If you prefer to force a precise placement, -ngl continues to work exactly as before; --fit applies only to parameters you haven’t set yourself. The behavior change concerns the default value, not the command.

#Which backend for your card

llama.cpp is compiled, or downloaded, for a given compute backend. The right choice depends solely on your hardware. The table covers the families in our database of 90 configurations.

Families derived from the QuelLLM hardware database (90 GPUs and chips) · 20/09/2026
Your hardwareRecommended backendNote
NVIDIA GTX 10 to RTX 50, desktop and laptopCUDAThe fastest and best-tested path.
AMD Radeon RX 7000 and RX 9000ROCm (HIP) or VulkanROCm is faster when it works. Vulkan installs without friction, including on Windows.
AMD Radeon RX 6000 and olderVulkanPartial or no ROCm support depending on the GPU.
Apple M1 to M5MetalEnabled by default in macOS binaries. All unified memory is usable.
Integrated Intel or AMD GPU, Intel ArcVulkan or SYCLReal-world gain over CPU-only operation on small models.
No GPUCPU (AVX2, AVX-512, NEON)Works everywhere. Target models with at most 8 billion parameters.

#Our measurements with llama.cpp

llama.cpp is our benchmark’s reference engine precisely because it runs identically on every platform. Here are the generation speeds measured on Llama 3.1 8B in Q4, with a 2,048-token context, a single request, and Flash Attention enabled.

Measured · QuelLLM benchmark, llama.cpp b4280, readings from 18/04/2026 · full table on the Benchmarks page
MachineMemoryLlama 3.1 8B Q4
RTX 509032 GB172 tok/s
RTX 409024 GB128 tok/s
RTX 407012 GB76 tok/s
Mac M3 Max64 GB unified64 tok/s
RTX 306012 GB44 tok/s
Ryzen 7 7700, CPU onlySystem RAM7.8 tok/s

Two takeaways. First, the gap between a RTX 3060 and a standalone processor is a factor of five to six: even an entry-level card changes the experience. Second, these figures would be practically identical in Ollama or LM Studio on the same machines, since the compute engine is the same.

#Stay with Ollama or switch to llama.cpp

Stay on Ollama or LM Studio if…Switch to llama.cpp if…
You want to chat with a model without reading documentation.Your model exceeds VRAM capacity and you want to fine-tune the GPU/CPU split.
You frequently switch models and appreciate the built-in catalog.You want to test an architecture released this week.
Your tools (Open WebUI, editor extensions) expect the Ollama API.You’re building a durable server and want to control every option, from context to KV cache.

#Get started in ten minutes

No need to compile anything to try it. Prebuilt binaries are published for Windows, macOS, and Linux with every release, and package managers take care of the rest. The command below downloads a small model from Hugging Face and opens the web interface at http://localhost:8080.

Install, then start the server
# macOS et Linux (Homebrew)
brew install llama.cpp

# Windows
winget install llama.cpp

# Télécharger un modèle et ouvrir l'interface web
llama-server -hf ggml-org/gemma-3-4b-it-GGUF -ngl 99 -c 8192

For a more capable model from the catalog, download the GGUF file of your choice and point to it with the -m option. Compiling from source is only useful for enabling a specific backend or tracking day-to-day development.

#Most common loading errors

Most issues with llama.cpp boil down to three causes: running out of memory, a GGUF file incompatible with the installed build, or a misspelled option. The command-line error message almost always indicates which one, provided you read it to the end instead of closing it and rerunning the command.

Read the message before changing a setting
Message or symptomLikely causeTry it
cudaMalloc failed: out of memoryThe model, with the requested context, exceeds the available VRAMReduce -c, let --fit adjust -ngl, or choose a lighter quantization
unknown model architectureThe GGUF file uses a newer architecture than the installed binaryUpdate to the latest published version; new architectures arrive in llama.cpp first
Very slow generation despite a recent GPUNot all layers are offloaded to the GPU, often because there is not enough free VRAM.Check the load output (the “offloaded” line) and close other applications using the card
error: invalid argumentAn option renamed or misspelled in the commandCompare with llama-server --help, which lists the current short and long aliases

The loading line displayed when the server starts is worth reading in full once: it shows the number of layers actually offloaded to the GPU, the effective context size, and the type of KV cache used. That's often faster than randomly adding options until it works. Also keep an eye on the installed version: llama.cpp releases nightly builds very frequently, and a fix for your card or model may already have been released without yet making it into the Homebrew or winget package you installed the previous week.

#FAQ

Is llama.cpp faster than Ollama?+
With identical settings, no: Ollama uses the same compute engine, so generation speed is very similar. The difference comes from the settings. llama.cpp lets you precisely choose the number of layers on the GPU, the context size, and cache quantization, which can sometimes keep a model in VRAM that Ollama would have partially offloaded to the CPU.
Do you need to know how to compile to use llama.cpp?+
No. Precompiled binaries are published with every release for Windows, macOS, and Linux, and the package is available through Homebrew (macOS, Linux) and winget (Windows): one installation command is all it takes, with no compiler involved. Building from source is only useful for enabling a specific backend missing from the official binary, applying a local patch, or following the development branch day by day instead of using a released version.
Can you use llama.cpp without a graphics card?+
Yes, that is actually its original use case: the project was designed to run a model on a simple processor, without a GPU or Python. On a recent processor, an 8B model in Q4 runs at around 8 tokens per second in our benchmark, which is still slow but readable for chatting or summarizing. Above 14 billion parameters, the wait becomes painful without hardware acceleration.
What’s the difference between llama.cpp and vLLM?+
llama.cpp targets a personal machine, regardless of hardware, with models quantized in GGUF. vLLM targets a GPU server handling many requests in parallel, with Hugging Face weights. The former maximizes what runs on your machine; the latter maximizes the throughput of a shared GPU.
And on Mac, llama.cpp or MLX?+
Both use the Apple GPU through unified memory. MLX, the library designed by Apple for its own chips, is often faster on recent models; llama.cpp offers its huge catalog of already-quantized GGUF files and finer controls for context and the KV cache. In practice, many Mac users keep both installed and choose based on the format available for the model they want.

Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.