Advanced 15 minllama.cpp

Compiling llama.cpp with CUDA: a step-by-step build guide pas

Ollama and LM Studio wrap llama.cpp — but are often still 1 or 2 versions behind, with conservative flags. Compiling llama.cpp yourself lets you recover 20 to 40% speed on a recent RTX and access features that have just been released. This guide covers Windows, Linux, and the CUDA build.

By Mohamed Meguedmi·Update 2026-03-05·Tested on Windows, macOS, and Linux

#Why compile it yourself

Preview features
New quantizations (IQ4_NL, Q2_K_S), new models (Mamba, DeepSeek), experimental flags.
Optimal performance
Build optimized for your specific CPU (AVX-512, AMX) and your CUDA compute capability (sm_89 for RTX 40, sm_90 for RTX 50).
Low-level tools
llama-bench for benchmarking, llama-cli for scripting, llama-server for a bare endpoint (without a UI or dependencies).
i
Who should compile?
If you're a developer, curious about performance, or working on Mamba/Jamba not supported by Ollama: yes. Otherwise, Ollama handles 95% of the work without compilation.

#Prerequisites

CUDA Toolkit 12.x
Available for download at developer.nvidia.com. Don’t confuse it with the drivers—the Toolkit includes nvcc, the compiler.
CMake 3.21+
apt install cmake on Linux, via chocolatey on Windows.
C++ compiler
gcc 11+ on Linux, MSVC 2022 Build Tools on Windows.
Git
For cloning and updating.

#1. Clone the repo

Terminal
git clone https://github.com/ggerganov/llama.cpp.git
cd llama.cpp

#2. Build with CUDA

Linux / macOS
cmake -B build -DGGML_CUDA=ON -DCMAKE_BUILD_TYPE=Release
cmake --build build --config Release -j $(nproc)
Windows (PowerShell)
cmake -B build -DGGML_CUDA=ON -DCMAKE_BUILD_TYPE=Release
cmake --build build --config Release -j

Compilation takes 3 to 10 minutes depending on your CPU. The binaries are placed in build/bin/ (Linux) or build/bin/Release/ (Windows).

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates
→
Build optimized for your GPU
Add -DCMAKE_CUDA_ARCHITECTURES=89 to target only RTX 40 (or 90 for RTX 50). The binary is smaller and the build is faster.

#3. Test

Downloading a GGUF
# Via Hugging Face CLI
pip install huggingface_hub
huggingface-cli download \
  TheBloke/Mistral-7B-Instruct-v0.2-GGUF \
  mistral-7b-instruct-v0.2.Q4_K_M.gguf \
  --local-dir ./models
Inference
./build/bin/llama-cli \
  -m ./models/mistral-7b-instruct-v0.2.Q4_K_M.gguf \
  -p "Explique ce qu'est un LLM en 3 phrases." \
  -n 256 \
  -ngl 99  # nombre de couches sur GPU
OpenAI-compatible HTTP server
./build/bin/llama-server \
  -m ./models/mistral-7b-instruct-v0.2.Q4_K_M.gguf \
  --port 8080 \
  -ngl 99 \
  -c 8192

#4. Optimization flags

-ngl N
Number of layers loaded on the GPU. Set it to 99 (or higher) to load everything if VRAM allows.
-fa
Flash Attention. +20-30% speed, reduced context VRAM usage. Always enable it if your GPU is Ampere+ (RTX 30xx and newer).
-c N
Context length. Don't use more than necessary: the KV cache eats VRAM.
-b N / -ub N
Batch size and micro-batch. 512/512 by default. 1024/512 to speed up prompt processing.
--no-mmap
Disables mmap. Worth testing if initial loading is slow on some SSDs.

#5. Stay up to date

Update
cd llama.cpp
git pull
cmake --build build --config Release -j

llama.cpp produces several commits per day. A weekly git pull is a good cadence. In case of a regression, use git checkout on a specific tag (e.g., b4400).

#Troubleshooting

nvcc not found
The CUDA Toolkit is not in PATH. On Linux: export PATH=/usr/local/cuda/bin:$PATH.
CUDA version mismatch
Drivers NVIDIA older than the Toolkit. Update the drivers (a higher version is required for each Toolkit).
Link error with cuBLAS
Incorrect LD_LIBRARY_PATH. Add /usr/local/cuda/lib64.
Build fails on MSVC
Open an “x64 Native Tools Command Prompt for VS 2022,” not PowerShell. Otherwise, some cmake tools cannot find cl.exe.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.