Intermediate 11 minInterfaces

KoboldCpp: install, configure (GGUF, ROCm, API)—and vs Ollama

KoboldCpp is a single binary that loads any GGUF model without installation or dependencies, with a web interface and API included. Built on llama.cpp, it shines where Ollama struggles: AMD cards via ROCm or Vulkan, fine-grained offloading control, and complete portability. This guide covers downloading, first launch, memory settings, and cases where KoboldCpp is a better replacement for Ollama.

By Thomas P.·Update 2026-09-05·Tested on Windows, macOS, and Linux
i
In brief
KoboldCpp is a single binary (a fork of llama.cpp) that loads any GGUF file without installation or dependencies. · It includes a web interface (KoboldAI Lite) and an OpenAI-compatible API in the same executable. · On AMD, it offers dedicated ROCm builds and a universal Vulkan backend, whereas Ollama is more limited. · Verdict: Ollama remains more convenient for a turnkey model catalog; KoboldCpp wins on portability and fine-grained GGUF tuning.

#Why KoboldCpp

KoboldCpp is a fork of llama.cpp packaged as a single standalone executable. You download a file, run it, and immediately get a web chat interface in your browser plus an HTTP API. No daemon to install, no Python dependencies to manage, and no proprietary model manager: point the binary to a GGUF file downloaded from Hugging Face, and that's it.

Compared with Ollama, the difference in philosophy is clear. Ollama manages a model catalog, a background daemon (on http://localhost:11434), and its own packaging format. KoboldCpp, by contrast, manages nothing: it directly runs the GGUF you provide. This makes it ideal when you want to test a specific quant downloaded manually, when you’re on a machine where you can’t install anything, or when your GPU is an AMD model that is poorly supported elsewhere.

Single binary
A single executable file, portable, with no installation or admin rights required.
Direct GGUF
Load any .gguf downloaded from Hugging Face, with no conversion.
Best-in-class AMD
Dedicated ROCm builds and a Vulkan backend that truly take advantage of Radeon cards.
Everything included
KoboldAI Lite web interface + API (native and OpenAI-compatible) in the same binary.
i
Roleplay origins
KoboldCpp comes from the KoboldAI ecosystem, which is heavily focused on writing and role-playing. The result: richer-than-average sampling and memory settings. But it remains an excellent general-purpose runner for chat, code, or RAG.

#Prerequisites

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

KoboldCpp runs on Windows, Linux, and macOS. It works on pure CPU, but a GPU greatly speeds up inference. The key point, as with any GGUF runner, is the available VRAM: it determines the model size and quantization you can load entirely onto the card.

System RAM
8 GB minimum for a small CPU model, 16 GB for comfortable use, 32 GB for offloading large models.
VRAM (Q4_K_M reference points)
3B ≈ 2 GB · 7B ≈ 5 GB · 14B ≈ 9 GB · 32B ≈ 19 GB · 70B ≈ 40 GB.
GPU NVIDIA
CUDA builds. RTX 3060 12GB (entry-level), 4070 12GB, 4080 16GB, 4090 24GB.
AMD GPU
ROCm builds (Radeon RX 6000/7000) or universal Vulkan backend.
A GGUF file
Download from Hugging Face (e.g., bartowski, one of the leading GGUF quant providers).
→
Which quantization to choose
Q4_K_M is the best quality-to-size compromise for most uses. Move up to Q5_K_M or Q8_0 if VRAM allows and you want more precision. The “Choosing Your Quantization” guide details the tradeoffs.

#Download the binary

Everything is done through the project’s GitHub releases (LostRuins/koboldcpp). Choose the binary that matches your OS and GPU—there is no installer, just a file you need to make executable.

  1. 01
    Open the releases
    Go to github.com/LostRuins/koboldcpp/releases and find the latest stable version.
  2. 02
    Choose the right file
    Windows NVIDIA: koboldcpp.exe (CUDA build). Windows without NVIDIA: koboldcpp_nocuda.exe (uses Vulkan/CLBlast). Linux: koboldcpp-linux-x64. AMD on Linux: the ROCm build koboldcpp-linux-x64-rocm (see the AMD section).
  3. 03
    Make executable (Linux/macOS)
    Download it, then grant the file execute permission before running it.
Terminal (Linux)
# Récupérer le binaire (adaptez l'URL à la dernière release)
wget https://github.com/LostRuins/koboldcpp/releases/latest/download/koboldcpp-linux-x64

# Le rendre exécutable
chmod +x koboldcpp-linux-x64

# Vérifier qu'il se lance
./koboldcpp-linux-x64 --help
i
Windows: double-click works
On Windows, launching the .exe without arguments opens a graphical configuration interface (the “launcher”). You choose the model, GPU backend, and context there with the mouse, without using the command line.

#Load a GGUF and chat

The core operation comes down to one command: pass the GGUF file path with --model. KoboldCpp starts a local server and opens—or shows you—the web interface URL, by default http://localhost:5001.

Terminal
# Lancer avec un modèle, en offloadant toutes les couches sur le GPU
./koboldcpp-linux-x64 \
  --model ./qwen3.5-9b-instruct-Q4_K_M.gguf \
  --gpulayers 999 \
  --contextsize 8192
--model
Path to the .gguf file to load.
--gpulayers
Number of layers offloaded to the GPU. 999 = everything that fits in VRAM.
--contextsize
Context window size in tokens (e.g., 4096, 8192, 16384).
--port
Listening port (5001 by default).
--host
Listening address. 0.0.0.0 to expose it on the local network.

Once launched, open http://localhost:5001 in your browser: you’ll land on KoboldAI Lite, a full chat interface with history, sampling settings, and modes (chat, instruct, writing). No other installation is required.


#AMD GPU: ROCm and Vulkan

This is where KoboldCpp stands out most from Ollama. There are two ways to accelerate on a Radeon, depending on your system and your card.

#The ROCm path (Linux, maximum performance)

ROCm is AMD's GPU computing stack, equivalent to CUDA. The project provides dedicated ROCm builds that deliver the best performance on Radeon RX 6000/7000. ROCm must be installed on the system, then you launch the ROCm binary exactly like the others.

Terminal (AMD ROCm)
# Build ROCm dédié
chmod +x koboldcpp-linux-x64-rocm

# Certaines cartes non officiellement supportées nécessitent de forcer
# la version d'architecture GPU (ex. RX 6700 XT -> gfx1030)
export HSA_OVERRIDE_GFX_VERSION=10.3.0

./koboldcpp-linux-x64-rocm \
  --model ./gemma-4-12b-it-Q4_K_M.gguf \
  --usecublas \
  --gpulayers 999 \
  --contextsize 8192
!
HSA_OVERRIDE_GFX_VERSION
Many consumer Radeon cards aren't on the official ROCm list and fail at startup. The HSA_OVERRIDE_GFX_VERSION variable forces a nearby compatible architecture (e.g., 10.3.0 for the RDNA2 generation). This is the setting that unlocks most cards.

#The Vulkan path (universal, simple)

If ROCm puts you off or is unavailable (Windows, an older card, or an iGPU), the Vulkan backend is an excellent alternative. It is vendor-agnostic: it runs on AMD, Intel, and NVIDIA without a specific compute stack, at the cost of a slight performance overhead compared with ROCm or CUDA.

Terminal (Vulkan)
./koboldcpp-linux-x64 \
  --model ./granite-4.2-8b-instruct-Q4_K_M.gguf \
  --usevulkan \
  --gpulayers 999 \
  --contextsize 8192
→
What to choose
On Linux with a recent Radeon and ROCm installed: use ROCm for speed. Everywhere else (Windows, iGPU, mixed hardware): Vulkan works out of the box. Compare tokens/sec throughput on your machine; the gap varies by model.

#Context and offloading

Two settings determine whether your model fits on the GPU and how quickly it responds: the number of offloaded layers and the context size. Calibrating them properly prevents spilling into RAM, which causes the speed to drop.

#Layer offloading (--gpulayers)

A model consists of layers. Each layer placed in VRAM is computed by the GPU; the layers that remain run on the CPU. --gpulayers 999 attempts to put everything on the GPU. If the card does not have enough VRAM, reduce this number: the model is then split between the GPU and CPU (partial offloading), which is slower but functional.

Everything fits in VRAM
--gpulayers 999, maximum speed, the entire model on the GPU.
Insufficient VRAM
Lower the value (e.g., 20, 30) until loading completes without saturating the card.
No GPU
--gpulayers 0, CPU-only: slow but works everywhere.

#Context window (--contextsize)

The context is the amount of text (in tokens) that the model keeps in memory: the system prompt, history, and question. The larger it is, the more VRAM the KV cache uses. Don’t increase the context beyond what you need: 4096 to 8192 is enough for everyday chat, while 16384+ is for analyzing long documents.

!
The bloated-context trap
Passing --contextsize 32768 “just in case” reserves a huge KV cache that can overflow VRAM on its own and force CPU offload. Result: the model fit at 8k but crawls at 32k. Adjust the context to your actual use case.

#Integrated API and web interface

KoboldCpp exposes two APIs on the same port (5001 by default): its own native KoboldAI API and an OpenAI-compatible API under /v1. The latter lets you connect KoboldCpp behind any tool that speaks the OpenAI protocol — exactly like an Ollama endpoint or llama-server.

Terminal (OpenAI API test)
curl http://localhost:5001/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "koboldcpp",
    "messages": [
      {"role": "user", "content": "Explique le offloading GPU en une phrase."}
    ]
  }'

On the Python side, OpenAI compatibility lets you reuse the official SDK by simply changing the base URL and using a dummy key.

Python (OpenAI SDK)
from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:5001/v1",
    api_key="koboldcpp",  # non vérifiée en local
)

resp = client.chat.completions.create(
    model="koboldcpp",
    messages=[{"role": "user", "content": "Bonjour en une phrase."}],
)
print(resp.choices[0].message.content)
Web interface
KoboldAI Lite on http://localhost:5001: chat, advanced sampling, personas, memory.
OpenAI API
/v1/chat/completions endpoint for connecting Open WebUI, scripts, or agents.
Native API
KoboldAI endpoints for fine-grained control over sampling and generation.
→
Combine with Open WebUI
Because the endpoint is OpenAI-compatible, you can use KoboldCpp as a backend behind Open WebUI: enter http://localhost:5001/v1 as an OpenAI connection in the interface settings.

#Troubleshooting

Loading fails on AMD
Card not recognized by ROCm: set HSA_OVERRIDE_GFX_VERSION to the closest architecture (e.g. 10.3.0), or switch to --usevulkan.
Very slow generation
The model spills over to the CPU. Lower --contextsize, reduce --gpulayers to avoid saturation, or drop down one quantization level (Q5 → Q4_K_M).
“out of memory” at startup
VRAM saturated by the model and KV cache. Reduce the context or offloading, or use a lighter quant.
Port already in use
Change it with --port (e.g. --port 5002) if 5001 is occupied.
Access from another machine
Add --host 0.0.0.0 to listen on the local network, and open the port in the firewall.

In short, KoboldCpp is the pragmatic choice when you want zero installation, direct control over the GGUF file and offloading, or simply want to finally make an AMD card work seriously. For a model catalog and turnkey system integration, Ollama remains more convenient; for portability and fine-tuning, KoboldCpp wins.


#Go further

These guides build on this one and cover adjacent components:

Ollama with AMD GPUs (ROCm)
The other AMD approach, on the Ollama side, for comparing ROCm across the two ecosystems.
Choose your quantization (Q4, Q5, Q8, FP16)
For balancing model size, VRAM, and quality before downloading a GGUF.
llama-server: a local OpenAI API with llama.cpp
The closest alternative, based on the same llama.cpp foundation, if you prioritize the API.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.