llama-server: local OpenAI API with llama.cpp
llama-server is the HTTP server built into llama.cpp. With one command (llama-server -m modele.gguf -ngl 99), it loads a GGUF and exposes an OpenAI-compatible API at http://localhost:8080, with a web interface included. Unlike Ollama, it provides direct control over GPU offloading (--n-gpu-layers), context, and batching, without a daemon or abstraction layer.
Ollama is convenient, but it hides everything from you: where your layers go, how the context is configured, and what is actually running on the GPU. llama-server, the HTTP server shipped with llama.cpp, does the opposite. A single command serves any GGUF file behind an OpenAI-compatible API, with a web interface and full control over offloading. This guide shows you how to launch it, connect your applications to it, and when it is a better choice than Ollama.
#Why llama-server?
Ollama, LM Studio, and Jan all rely on the same underlying engine: llama.cpp. This engine includes its own HTTP server, llama-server, which needs none of these layers. Point it to a GGUF file, and you get an API and a web interface. Nothing more.
The benefit is not cosmetic. Where Ollama decides for you how many layers to send to the GPU, the context size, and how to split the model, llama-server exposes every parameter on the command line. You see and control what happens. This is the “manual” mode of local AI—more verbose, but without a black box.
- OpenAI-compatible API
- Endpoints /v1/chat/completions, /v1/completions, /v1/models, /v1/embeddings. Any OpenAI client can connect to them without modification.
- Offloading control
- --n-gpu-layers precisely sets how many layers are placed in VRAM. Essential when the model exceeds your card's capacity.
- Built-in web interface
- A chat served directly from the server root, without installing Open WebUI or Docker.
- Zero heavy dependencies
- A single binary (a few dozen MB). No Python, no container, and no mandatory system service.
- Batching and parallelism
- Continuous batching enabled by default, with multiple simultaneous requests through slots.
#Prerequisites
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
- A GGUF file
- The llama.cpp format. Available from Hugging Face, or downloaded directly by llama-server via -hf (see below).
- VRAM or RAM
- In Q4_K_M, allow ≈2 GB for a 3B, ≈5 GB for a 7B, ≈9 GB for a 14B, ≈19 GB for a 32B, and ≈40 GB for a 70B.
- A GPU (strongly recommended)
- RTX 3060 12 GB to get started, RTX 4070/4080 for the mid-range, RTX 4090 24 GB or a Mac M4 Pro for large models. The CPU alone works, but slowly.
- A terminal
- llama-server is controlled from the command line. Nothing insurmountable, but it is not a double-click app like LM Studio.
#1. Obtain llama-server
Three paths, from fastest to highest-performing. On macOS, Homebrew installs the binary with one command:
On Windows and Linux, the simplest option is to download a precompiled binary from the official llama.cpp GitHub releases (choose the variant matching your hardware: CUDA for NVIDIA, Vulkan for a generic GPU, or CPU).
For maximum tokens per second, compile from source with your GPU's backend. Example for NVIDIA with CUDA:
#2. Run a GGUF with one command
The minimal command selects a model and starts the server. Here, a Qwen 3.5 9B in Q4_K_M (the reference 8 GB choice in 2026, with 256k context) with all layers offloaded to the GPU:
- -m
- Path to the GGUF file to serve.
- -ngl 99
- Number of layers offloaded to the GPU. 99 means “all” (the model has fewer; the excess is ignored without error).
- -c 8192
- Context size in tokens. The default is often 4096; adjust it according to your needs and VRAM.
Don't have the file on hand? llama-server can download it directly from Hugging Face and cache it, like built-in ollama pull mais:
Once started, the server listens by default on http://127.0.0.1:8080. Verify that it is alive:
#3. The OpenAI-compatible API: connect any app
That’s llama-server’s killer feature. It speaks the OpenAI protocol, so any tool designed for the OpenAI API works by simply changing the base URL. A direct chat call with curl:
The model field is unrestricted: llama-server serves only one model at a time and largely ignores this value. With the OpenAI Python SDK, simply redirect base_url to your server. The API key can be any string if you have not set --api-key:
- /v1/chat/completions
- Conversation mode, with automatic application of the model's chat template.
- /v1/completions
- Raw text completion, without role formatting.
- /v1/models
- Lists the loaded model—useful for clients that first query the available models.
- /v1/embeddings
- Generates embeddings if the server is started with --embedding (useful for a custom RAG setup).
#4. n-gpu-layers: fine-grained offloading that Ollama hides
A model is a stack of layers. Each layer sent to VRAM is computed by the GPU, very quickly; those that remain in RAM are computed by the CPU, slowly. --n-gpu-layers (or -ngl) determines how many layers go to the GPU. It's the most important setting for speed.
- -ngl 99
- All about the GPU. Aim for this if the model fits entirely in VRAM. Maximum speed.
- -ngl 20
- Partial offloading: 20 layers on the GPU, the rest on the CPU. The compromise when the model exceeds VRAM.
- -ngl 0
- All on the CPU. Slow, but lets you run a much larger model than your card can handle.
The strategy: raise -ngl as high as possible without saturating VRAM. A 14B in Q4 (≈9 GB) fits entirely on a RTX 3060 12 GB with -ngl 99. A Qwen 3.8 27B in Q4 (≈18 GB) does not fit; on that same card, offload partially—for example, -ngl 40—and accept a slowdown.
To monitor what actually fits in VRAM during loading, keep an eye on nvidia-smi in another window:
#5. Included web interface
You don’t need Open WebUI or Docker to chat: llama-server serves a chat interface directly at its root. Simply open the server address in a browser.
You get a complete chat interface: conversation history, temperature and sampling-parameter controls, a system-prompt system, and Markdown rendering. That's enough for everyday personal use without installing any additional layer.
#When to prefer llama-server over Ollama (and when to stay)
llama-server and Ollama run the same engine. The choice is about control versus convenience.
- Choose llama-server
- When you want to fine-tune offloading, test a specific GGUF from a particular quantizer, avoid a permanent daemon, or deploy a single binary without dependencies on a server.
- Choose llama-server
- When a model exceeds your VRAM: direct control of -ngl and memory options makes the difference between “unplayable” and “slow but functional.”
- Stay on Ollama
- When you want to switch between multiple models on the fly without restarting processes, manage a library with ollama pull/list, or automatically load and unload models based on demand.
- Stay on Ollama
- When multiple applications target different models on the same port 11434: Ollama routes and swaps for you, whereas llama-server runs a single model per process.
#Troubleshooting
- “CUDA out of memory” while loading
- Your -ngl is too high for the VRAM. Lower it (partial offloading), reduce -c, or switch to a lighter quantization (Q4_K_M instead of Q5/Q8).
- The GPU is not being used
- The binary may be the CPU variant. Check that you are using a CUDA/Metal/Vulkan build and that -ngl is greater than 0. nvidia-smi should show occupied VRAM.
- Inconsistent responses or visible tags
- The chat template was not applied. Rerun with --jinja to use the template embedded in the GGUF.
- The client app cannot find the model
- Some clients query /v1/models first. Give it an alias with -a and enter that exact name in your application’s model field.
- Truncated context / cut-off responses
- -c is too small. Increase the context size, bearing in mind that a large context uses more VRAM.
#Go further
llama-server gets its full value from a well-compiled llama.cpp and a well-chosen GGUF. These site guides round out the setup:
- Compile llama.cpp with CUDA
- For an optimized NVIDIA binary and maximum tokens per second on your card.
- Q4, Q5, Q8: which quantization should you choose
- To weigh quality, speed, and VRAM before downloading a GGUF.
- llama.cpp vs vLLM vs Exllama
- To compare llama-server with other inference engines based on your throughput needs.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.