Advanced 14 minllama.cpp

llama-server: local OpenAI API with llama.cpp

Direct response

llama-server is the HTTP server built into llama.cpp. With one command (llama-server -m modele.gguf -ngl 99), it loads a GGUF and exposes an OpenAI-compatible API at http://localhost:8080, with a web interface included. Unlike Ollama, it provides direct control over GPU offloading (--n-gpu-layers), context, and batching, without a daemon or abstraction layer.

Ollama is convenient, but it hides everything from you: where your layers go, how the context is configured, and what is actually running on the GPU. llama-server, the HTTP server shipped with llama.cpp, does the opposite. A single command serves any GGUF file behind an OpenAI-compatible API, with a web interface and full control over offloading. This guide shows you how to launch it, connect your applications to it, and when it is a better choice than Ollama.

By Mohamed Meguedmi·Update 2026-08-27·Tested on Windows, macOS, and Linux

#Why llama-server?

Ollama, LM Studio, and Jan all rely on the same underlying engine: llama.cpp. This engine includes its own HTTP server, llama-server, which needs none of these layers. Point it to a GGUF file, and you get an API and a web interface. Nothing more.

The benefit is not cosmetic. Where Ollama decides for you how many layers to send to the GPU, the context size, and how to split the model, llama-server exposes every parameter on the command line. You see and control what happens. This is the “manual” mode of local AI—more verbose, but without a black box.

i
In two words
llama-server is the HTTP binary for llama.cpp. One process, one GGUF model, an OpenAI-compatible API on port 8080, and an included web interface. No background daemon and no model library managed for you.
OpenAI-compatible API
Endpoints /v1/chat/completions, /v1/completions, /v1/models, /v1/embeddings. Any OpenAI client can connect to them without modification.
Offloading control
--n-gpu-layers precisely sets how many layers are placed in VRAM. Essential when the model exceeds your card's capacity.
Built-in web interface
A chat served directly from the server root, without installing Open WebUI or Docker.
Zero heavy dependencies
A single binary (a few dozen MB). No Python, no container, and no mandatory system service.
Batching and parallelism
Continuous batching enabled by default, with multiple simultaneous requests through slots.

#Prerequisites

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates
A GGUF file
The llama.cpp format. Available from Hugging Face, or downloaded directly by llama-server via -hf (see below).
VRAM or RAM
In Q4_K_M, allow ≈2 GB for a 3B, ≈5 GB for a 7B, ≈9 GB for a 14B, ≈19 GB for a 32B, and ≈40 GB for a 70B.
A GPU (strongly recommended)
RTX 3060 12 GB to get started, RTX 4070/4080 for the mid-range, RTX 4090 24 GB or a Mac M4 Pro for large models. The CPU alone works, but slowly.
A terminal
llama-server is controlled from the command line. Nothing insurmountable, but it is not a double-click app like LM Studio.
→
Are you starting from Ollama?
Models downloaded by Ollama are already GGUF files, stored in its blobs directory under a hashed name. Simpler: download the desired GGUF from Hugging Face or let llama-server retrieve it with -hf.

#1. Obtain llama-server

Three paths, from fastest to highest-performing. On macOS, Homebrew installs the binary with one command:

macOS — Homebrew
brew install llama.cpp

# le binaire s'appelle llama-server
llama-server --version

On Windows and Linux, the simplest option is to download a precompiled binary from the official llama.cpp GitHub releases (choose the variant matching your hardware: CUDA for NVIDIA, Vulkan for a generic GPU, or CPU).

Official releases
https://github.com/ggml-org/llama.cpp/releases

For maximum tokens per second, compile from source with your GPU's backend. Example for NVIDIA with CUDA:

Compile with CUDA
git clone https://github.com/ggml-org/llama.cpp
cd llama.cpp
cmake -B build -DGGML_CUDA=ON
cmake --build build --config Release -j

# le binaire se trouve dans build/bin/
./build/bin/llama-server --version
i
The binary was renamed
Historically, this server was llama.cpp’s “server” example. Since the tools were reorganized, it is called llama-server. If an old tutorial refers to ./server, it is the same program.

#2. Run a GGUF with one command

The minimal command selects a model and starts the server. Here, a Qwen 3.5 9B in Q4_K_M (the reference 8 GB choice in 2026, with 256k context) with all layers offloaded to the GPU:

Terminal
llama-server -m ./qwen3.5-9b-instruct-q4_k_m.gguf -ngl 99 -c 8192
-m
Path to the GGUF file to serve.
-ngl 99
Number of layers offloaded to the GPU. 99 means “all” (the model has fewer; the excess is ignored without error).
-c 8192
Context size in tokens. The default is often 4096; adjust it according to your needs and VRAM.

Don't have the file on hand? llama-server can download it directly from Hugging Face and cache it, like built-in ollama pull mais:

Download from Hugging Face
llama-server -hf bartowski/Qwen3.5-9B-Instruct-GGUF:Q4_K_M -ngl 99 -c 8192

Once started, the server listens by default on http://127.0.0.1:8080. Verify that it is alive:

Health check
curl http://localhost:8080/health
# {"status":"ok"}
!
Exposing it to the network = risk
By default, llama-server listens only on localhost. To make it accessible from other machines, add --host 0.0.0.0 — but then the API is open to anyone on the network. Protect it with --api-key and, ideally, an HTTPS reverse proxy.

#3. The OpenAI-compatible API: connect any app

That’s llama-server’s killer feature. It speaks the OpenAI protocol, so any tool designed for the OpenAI API works by simply changing the base URL. A direct chat call with curl:

/v1/chat/completions call
curl http://localhost:8080/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "local",
    "messages": [
      {"role": "system", "content": "Tu réponds en français, de façon concise."},
      {"role": "user", "content": "Explique le format GGUF en une phrase."}
    ],
    "temperature": 0.7
  }'

The model field is unrestricted: llama-server serves only one model at a time and largely ignores this value. With the OpenAI Python SDK, simply redirect base_url to your server. The API key can be any string if you have not set --api-key:

OpenAI Python client
from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:8080/v1",
    api_key="sk-no-key-required",
)

resp = client.chat.completions.create(
    model="local",
    messages=[
        {"role": "user", "content": "Donne-moi trois idées de noms pour un projet open source."}
    ],
)
print(resp.choices[0].message.content)
/v1/chat/completions
Conversation mode, with automatic application of the model's chat template.
/v1/completions
Raw text completion, without role formatting.
/v1/models
Lists the loaded model—useful for clients that first query the available models.
/v1/embeddings
Generates embeddings if the server is started with --embedding (useful for a custom RAG setup).
→
The chat template matters
For the model to format the conversation correctly, add --jinja at launch: llama-server then applies the chat template embedded in the GGUF. Without it, some newer models respond incorrectly.

#4. n-gpu-layers: fine-grained offloading that Ollama hides

A model is a stack of layers. Each layer sent to VRAM is computed by the GPU, very quickly; those that remain in RAM are computed by the CPU, slowly. --n-gpu-layers (or -ngl) determines how many layers go to the GPU. It's the most important setting for speed.

-ngl 99
All about the GPU. Aim for this if the model fits entirely in VRAM. Maximum speed.
-ngl 20
Partial offloading: 20 layers on the GPU, the rest on the CPU. The compromise when the model exceeds VRAM.
-ngl 0
All on the CPU. Slow, but lets you run a much larger model than your card can handle.

The strategy: raise -ngl as high as possible without saturating VRAM. A 14B in Q4 (≈9 GB) fits entirely on a RTX 3060 12 GB with -ngl 99. A Qwen 3.8 27B in Q4 (≈18 GB) does not fit; on that same card, offload partially—for example, -ngl 40—and accept a slowdown.

Partial offloading of a large model
# Qwen 3.8 27B Q4 (~18 Go) sur un GPU 12 Go : une partie sur GPU, le reste sur CPU
llama-server -m ./qwen3.8-27b-instruct-q4_k_m.gguf -ngl 40 -c 4096

To monitor what actually fits in VRAM during loading, keep an eye on nvidia-smi in another window:

VRAM monitoring
nvidia-smi -l 1
i
Flash attention and batching
Add --flash-attn to reduce context memory usage on compatible GPUs. Continuous batching (-cb) is enabled by default: multiple concurrent requests are handled efficiently through parallel slots (--parallel N).

#5. Included web interface

You don’t need Open WebUI or Docker to chat: llama-server serves a chat interface directly at its root. Simply open the server address in a browser.

Web interface
http://localhost:8080

You get a complete chat interface: conversation history, temperature and sampling-parameter controls, a system-prompt system, and Markdown rendering. That's enough for everyday personal use without installing any additional layer.

→
Name the displayed model
Use -a (or --alias) to give the model a readable name, which is reused in /v1/models and in the interface: llama-server -m modele.gguf -a qwen3.5-9b -ngl 99.

#When to prefer llama-server over Ollama (and when to stay)

llama-server and Ollama run the same engine. The choice is about control versus convenience.

Choose llama-server
When you want to fine-tune offloading, test a specific GGUF from a particular quantizer, avoid a permanent daemon, or deploy a single binary without dependencies on a server.
Choose llama-server
When a model exceeds your VRAM: direct control of -ngl and memory options makes the difference between “unplayable” and “slow but functional.”
Stay on Ollama
When you want to switch between multiple models on the fly without restarting processes, manage a library with ollama pull/list, or automatically load and unload models based on demand.
Stay on Ollama
When multiple applications target different models on the same port 11434: Ollama routes and swaps for you, whereas llama-server runs a single model per process.
i
Both coexist
You don't have to choose. Many people keep Ollama for everyday convenience and run a dedicated llama-server for a specific model to serve in production or for an offloading configuration that Ollama doesn't allow.

#Troubleshooting

“CUDA out of memory” while loading
Your -ngl is too high for the VRAM. Lower it (partial offloading), reduce -c, or switch to a lighter quantization (Q4_K_M instead of Q5/Q8).
The GPU is not being used
The binary may be the CPU variant. Check that you are using a CUDA/Metal/Vulkan build and that -ngl is greater than 0. nvidia-smi should show occupied VRAM.
Inconsistent responses or visible tags
The chat template was not applied. Rerun with --jinja to use the template embedded in the GGUF.
The client app cannot find the model
Some clients query /v1/models first. Give it an alias with -a and enter that exact name in your application’s model field.
Truncated context / cut-off responses
-c is too small. Increase the context size, bearing in mind that a large context uses more VRAM.

#Go further

llama-server gets its full value from a well-compiled llama.cpp and a well-chosen GGUF. These site guides round out the setup:

Compile llama.cpp with CUDA
For an optimized NVIDIA binary and maximum tokens per second on your card.
Q4, Q5, Q8: which quantization should you choose
To weigh quality, speed, and VRAM before downloading a GGUF.
llama.cpp vs vLLM vs Exllama
To compare llama-server with other inference engines based on your throughput needs.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.