KoboldCpp: install, configure (GGUF, ROCm, API)—and vs Ollama
KoboldCpp is a single binary that loads any GGUF model without installation or dependencies, with a web interface and API included. Built on llama.cpp, it shines where Ollama struggles: AMD cards via ROCm or Vulkan, fine-grained offloading control, and complete portability. This guide covers downloading, first launch, memory settings, and cases where KoboldCpp is a better replacement for Ollama.
#Why KoboldCpp
KoboldCpp is a fork of llama.cpp packaged as a single standalone executable. You download a file, run it, and immediately get a web chat interface in your browser plus an HTTP API. No daemon to install, no Python dependencies to manage, and no proprietary model manager: point the binary to a GGUF file downloaded from Hugging Face, and that's it.
Compared with Ollama, the difference in philosophy is clear. Ollama manages a model catalog, a background daemon (on http://localhost:11434), and its own packaging format. KoboldCpp, by contrast, manages nothing: it directly runs the GGUF you provide. This makes it ideal when you want to test a specific quant downloaded manually, when you’re on a machine where you can’t install anything, or when your GPU is an AMD model that is poorly supported elsewhere.
- Single binary
- A single executable file, portable, with no installation or admin rights required.
- Direct GGUF
- Load any .gguf downloaded from Hugging Face, with no conversion.
- Best-in-class AMD
- Dedicated ROCm builds and a Vulkan backend that truly take advantage of Radeon cards.
- Everything included
- KoboldAI Lite web interface + API (native and OpenAI-compatible) in the same binary.
#Prerequisites
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
KoboldCpp runs on Windows, Linux, and macOS. It works on pure CPU, but a GPU greatly speeds up inference. The key point, as with any GGUF runner, is the available VRAM: it determines the model size and quantization you can load entirely onto the card.
- System RAM
- 8 GB minimum for a small CPU model, 16 GB for comfortable use, 32 GB for offloading large models.
- VRAM (Q4_K_M reference points)
- 3B ≈ 2 GB · 7B ≈ 5 GB · 14B ≈ 9 GB · 32B ≈ 19 GB · 70B ≈ 40 GB.
- GPU NVIDIA
- CUDA builds. RTX 3060 12GB (entry-level), 4070 12GB, 4080 16GB, 4090 24GB.
- AMD GPU
- ROCm builds (Radeon RX 6000/7000) or universal Vulkan backend.
- A GGUF file
- Download from Hugging Face (e.g., bartowski, one of the leading GGUF quant providers).
#Download the binary
Everything is done through the project’s GitHub releases (LostRuins/koboldcpp). Choose the binary that matches your OS and GPU—there is no installer, just a file you need to make executable.
- 01Open the releasesGo to github.com/LostRuins/koboldcpp/releases and find the latest stable version.
- 02Choose the right fileWindows NVIDIA: koboldcpp.exe (CUDA build). Windows without NVIDIA: koboldcpp_nocuda.exe (uses Vulkan/CLBlast). Linux: koboldcpp-linux-x64. AMD on Linux: the ROCm build koboldcpp-linux-x64-rocm (see the AMD section).
- 03Make executable (Linux/macOS)Download it, then grant the file execute permission before running it.
#Load a GGUF and chat
The core operation comes down to one command: pass the GGUF file path with --model. KoboldCpp starts a local server and opens—or shows you—the web interface URL, by default http://localhost:5001.
- --model
- Path to the .gguf file to load.
- --gpulayers
- Number of layers offloaded to the GPU. 999 = everything that fits in VRAM.
- --contextsize
- Context window size in tokens (e.g., 4096, 8192, 16384).
- --port
- Listening port (5001 by default).
- --host
- Listening address. 0.0.0.0 to expose it on the local network.
Once launched, open http://localhost:5001 in your browser: you’ll land on KoboldAI Lite, a full chat interface with history, sampling settings, and modes (chat, instruct, writing). No other installation is required.
#AMD GPU: ROCm and Vulkan
This is where KoboldCpp stands out most from Ollama. There are two ways to accelerate on a Radeon, depending on your system and your card.
#The ROCm path (Linux, maximum performance)
ROCm is AMD's GPU computing stack, equivalent to CUDA. The project provides dedicated ROCm builds that deliver the best performance on Radeon RX 6000/7000. ROCm must be installed on the system, then you launch the ROCm binary exactly like the others.
#The Vulkan path (universal, simple)
If ROCm puts you off or is unavailable (Windows, an older card, or an iGPU), the Vulkan backend is an excellent alternative. It is vendor-agnostic: it runs on AMD, Intel, and NVIDIA without a specific compute stack, at the cost of a slight performance overhead compared with ROCm or CUDA.
#Context and offloading
Two settings determine whether your model fits on the GPU and how quickly it responds: the number of offloaded layers and the context size. Calibrating them properly prevents spilling into RAM, which causes the speed to drop.
#Layer offloading (--gpulayers)
A model consists of layers. Each layer placed in VRAM is computed by the GPU; the layers that remain run on the CPU. --gpulayers 999 attempts to put everything on the GPU. If the card does not have enough VRAM, reduce this number: the model is then split between the GPU and CPU (partial offloading), which is slower but functional.
- Everything fits in VRAM
- --gpulayers 999, maximum speed, the entire model on the GPU.
- Insufficient VRAM
- Lower the value (e.g., 20, 30) until loading completes without saturating the card.
- No GPU
- --gpulayers 0, CPU-only: slow but works everywhere.
#Context window (--contextsize)
The context is the amount of text (in tokens) that the model keeps in memory: the system prompt, history, and question. The larger it is, the more VRAM the KV cache uses. Don’t increase the context beyond what you need: 4096 to 8192 is enough for everyday chat, while 16384+ is for analyzing long documents.
#Integrated API and web interface
KoboldCpp exposes two APIs on the same port (5001 by default): its own native KoboldAI API and an OpenAI-compatible API under /v1. The latter lets you connect KoboldCpp behind any tool that speaks the OpenAI protocol — exactly like an Ollama endpoint or llama-server.
On the Python side, OpenAI compatibility lets you reuse the official SDK by simply changing the base URL and using a dummy key.
- Web interface
- KoboldAI Lite on http://localhost:5001: chat, advanced sampling, personas, memory.
- OpenAI API
- /v1/chat/completions endpoint for connecting Open WebUI, scripts, or agents.
- Native API
- KoboldAI endpoints for fine-grained control over sampling and generation.
#Troubleshooting
- Loading fails on AMD
- Card not recognized by ROCm: set HSA_OVERRIDE_GFX_VERSION to the closest architecture (e.g. 10.3.0), or switch to --usevulkan.
- Very slow generation
- The model spills over to the CPU. Lower --contextsize, reduce --gpulayers to avoid saturation, or drop down one quantization level (Q5 → Q4_K_M).
- “out of memory” at startup
- VRAM saturated by the model and KV cache. Reduce the context or offloading, or use a lighter quant.
- Port already in use
- Change it with --port (e.g. --port 5002) if 5001 is occupied.
- Access from another machine
- Add --host 0.0.0.0 to listen on the local network, and open the port in the firewall.
In short, KoboldCpp is the pragmatic choice when you want zero installation, direct control over the GGUF file and offloading, or simply want to finally make an AMD card work seriously. For a model catalog and turnkey system integration, Ollama remains more convenient; for portability and fine-tuning, KoboldCpp wins.
#Go further
These guides build on this one and cover adjacent components:
- Ollama with AMD GPUs (ROCm)
- The other AMD approach, on the Ollama side, for comparing ROCm across the two ecosystems.
- Choose your quantization (Q4, Q5, Q8, FP16)
- For balancing model size, VRAM, and quality before downloading a GGUF.
- llama-server: a local OpenAI API with llama.cpp
- The closest alternative, based on the same llama.cpp foundation, if you prioritize the API.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.