Advanced 12 minllama.cpp

llama.cpp with Vulkan (GPU universel)

Vulkan is the cross-vendor graphics API: it works on NVIDIA, AMD, Intel Arc, and even iGPUs. Compiling llama.cpp with the Vulkan backend lets you completely avoid CUDA/ROCm headaches and have a solution that works everywhere. The tradeoff is 10 to 20% lower performance than a native backend—often a good compromise for heterogeneous setups.

By Mohamed Meguedmi·Update 2026-08-27·Tested on Windows, macOS, and Linux

#Why Vulkan

Vendor-agnostic
The same binary runs on GeForce, Radeon, Arc, and Iris. Convenient for distributing a tool or testing a model before buying a card.
Simple installation
Vulkan drivers are included with standard graphics drivers. No multi-GB Toolkit to install.
Older cards
A GTX 1060 6 GB or a RX 580 works through Vulkan, whereas CUDA has been abandoned and ROCm is unsupported.
Intel Arc
The Arc A580/A750/A770 are properly supported through Vulkan, whereas oneAPI is more unreliable.

#Compatible GPUs

NVIDIA
Any Maxwell+ card (GTX 900 and newer). Vulkan is supported in all recent NVIDIA drivers.
AMD
Polaris (RX 400/500), Vega, Navi (RX 5000/6000/7000/9000). Mesa on Linux, official drivers on Windows.
Intel Arc
Alchemist (A380-A770) and Battlemage. Excellent performance per dollar for local AI.
Intel iGPU
Xe (Tiger Lake+), Xe2 (Lunar Lake, Arrow Lake). Useful for small models on laptops without a dedicated GPU.

#Prerequisites

Ubuntu / Debian
sudo apt update
sudo apt install libvulkan-dev vulkan-tools glslang-tools \
  cmake build-essential git
Fedora
sudo dnf install vulkan-headers vulkan-tools \
  glslang-devel cmake gcc-c++ git
Windows
# Installer le SDK Vulkan LunarG (fournit les headers et shaderc)
winget install KhronosGroup.VulkanSDK
Verify that Vulkan detects the GPU
vulkaninfo | grep -i "deviceName"
# Doit afficher votre GPU

#1. Build

Terminal
git clone https://github.com/ggerganov/llama.cpp.git
cd llama.cpp

cmake -B build -DGGML_VULKAN=ON -DCMAKE_BUILD_TYPE=Release
cmake --build build --config Release -j

Vulkan shaders are compiled during the build. Allow 3-8 minutes depending on the CPU.

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

#2. Test

Run
./build/bin/llama-cli \
  -m ./models/qwen3.5-9b-q4_k_m.gguf \
  -p "Bonjour" -n 128 -ngl 99

At startup, llama.cpp displays the GPU detected through Vulkan: ggml_vulkan: Found 1 Vulkan devices: ... Intel(R) Arc(TM) A770. If you have multiple GPUs, use -mg N (main GPU index) to choose.

#3. Performance vs. CUDA / ROCm

Tokens per second on Qwen 3.5 9B Q4_K_M, with Flash Attention enabled (rough figures):

RTX 4070 — native CUDA
~82 tok/s (reference).
RTX 4070 — Vulkan
~70 tok/s (-15%).
RX 7800 XT — ROCm
~56 tok/s.
RX 7800 XT — Vulkan
~48 tok/s (-15%). Sometimes the fastest option when ROCm blocks.
Arc A770 16 GB — Vulkan
~43 tok/s. Best value for money.
Intel Xe iGPU (Lunar Lake)
~8–13 tok/s. Usable for a 2–3B model, but struggles with a 9B model.

#When to choose Vulkan

You have an Intel Arc card
Vulkan > oneAPI in 2026. Stable performance, no exotic installation.
ROCm refuses to cooperate
AMD card not officially supported, override crashes: Vulkan is the fallback that works.
Multi-vendor setup
A machine with GeForce + Intel Arc, a variable Thunderbolt+eGPU laptop. Vulkan tolerates everything.
Tool distribution
You're coding an app that must run on unknown machines: the Vulkan binary is the most compatible.
i
Not for Mac
On macOS, Vulkan runs through MoltenVK (a translation layer for Metal). Use llama.cpp's native Metal backend instead — it's faster.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.