Compiling llama.cpp with CUDA: a step-by-step build guide pas
Ollama and LM Studio wrap llama.cpp — but are often still 1 or 2 versions behind, with conservative flags. Compiling llama.cpp yourself lets you recover 20 to 40% speed on a recent RTX and access features that have just been released. This guide covers Windows, Linux, and the CUDA build.
#Why compile it yourself
- Preview features
- New quantizations (IQ4_NL, Q2_K_S), new models (Mamba, DeepSeek), experimental flags.
- Optimal performance
- Build optimized for your specific CPU (AVX-512, AMX) and your CUDA compute capability (sm_89 for RTX 40, sm_90 for RTX 50).
- Low-level tools
- llama-bench for benchmarking, llama-cli for scripting, llama-server for a bare endpoint (without a UI or dependencies).
#Prerequisites
- CUDA Toolkit 12.x
- Available for download at developer.nvidia.com. Don’t confuse it with the drivers—the Toolkit includes nvcc, the compiler.
- CMake 3.21+
- apt install cmake on Linux, via chocolatey on Windows.
- C++ compiler
- gcc 11+ on Linux, MSVC 2022 Build Tools on Windows.
- Git
- For cloning and updating.
#1. Clone the repo
#2. Build with CUDA
Compilation takes 3 to 10 minutes depending on your CPU. The binaries are placed in build/bin/ (Linux) or build/bin/Release/ (Windows).
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
#3. Test
#4. Optimization flags
- -ngl N
- Number of layers loaded on the GPU. Set it to 99 (or higher) to load everything if VRAM allows.
- -fa
- Flash Attention. +20-30% speed, reduced context VRAM usage. Always enable it if your GPU is Ampere+ (RTX 30xx and newer).
- -c N
- Context length. Don't use more than necessary: the KV cache eats VRAM.
- -b N / -ub N
- Batch size and micro-batch. 512/512 by default. 1024/512 to speed up prompt processing.
- --no-mmap
- Disables mmap. Worth testing if initial loading is slow on some SSDs.
#5. Stay up to date
llama.cpp produces several commits per day. A weekly git pull is a good cadence. In case of a regression, use git checkout on a specific tag (e.g., b4400).
#Troubleshooting
- nvcc not found
- The CUDA Toolkit is not in PATH. On Linux: export PATH=/usr/local/cuda/bin:$PATH.
- CUDA version mismatch
- Drivers NVIDIA older than the Toolkit. Update the drivers (a higher version is required for each Toolkit).
- Link error with cuBLAS
- Incorrect LD_LIBRARY_PATH. Add /usr/local/cuda/lib64.
- Build fails on MSVC
- Open an “x64 Native Tools Command Prompt for VS 2022,” not PowerShell. Otherwise, some cmake tools cannot find cl.exe.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.