llama.cpp with Vulkan (GPU universel)
Vulkan is the cross-vendor graphics API: it works on NVIDIA, AMD, Intel Arc, and even iGPUs. Compiling llama.cpp with the Vulkan backend lets you completely avoid CUDA/ROCm headaches and have a solution that works everywhere. The tradeoff is 10 to 20% lower performance than a native backend—often a good compromise for heterogeneous setups.
#Why Vulkan
- Vendor-agnostic
- The same binary runs on GeForce, Radeon, Arc, and Iris. Convenient for distributing a tool or testing a model before buying a card.
- Simple installation
- Vulkan drivers are included with standard graphics drivers. No multi-GB Toolkit to install.
- Older cards
- A GTX 1060 6 GB or a RX 580 works through Vulkan, whereas CUDA has been abandoned and ROCm is unsupported.
- Intel Arc
- The Arc A580/A750/A770 are properly supported through Vulkan, whereas oneAPI is more unreliable.
#Compatible GPUs
- NVIDIA
- Any Maxwell+ card (GTX 900 and newer). Vulkan is supported in all recent NVIDIA drivers.
- AMD
- Polaris (RX 400/500), Vega, Navi (RX 5000/6000/7000/9000). Mesa on Linux, official drivers on Windows.
- Intel Arc
- Alchemist (A380-A770) and Battlemage. Excellent performance per dollar for local AI.
- Intel iGPU
- Xe (Tiger Lake+), Xe2 (Lunar Lake, Arrow Lake). Useful for small models on laptops without a dedicated GPU.
#Prerequisites
#1. Build
Vulkan shaders are compiled during the build. Allow 3-8 minutes depending on the CPU.
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
#2. Test
At startup, llama.cpp displays the GPU detected through Vulkan: ggml_vulkan: Found 1 Vulkan devices: ... Intel(R) Arc(TM) A770. If you have multiple GPUs, use -mg N (main GPU index) to choose.
#3. Performance vs. CUDA / ROCm
Tokens per second on Qwen 3.5 9B Q4_K_M, with Flash Attention enabled (rough figures):
- RTX 4070 — native CUDA
- ~82 tok/s (reference).
- RTX 4070 — Vulkan
- ~70 tok/s (-15%).
- RX 7800 XT — ROCm
- ~56 tok/s.
- RX 7800 XT — Vulkan
- ~48 tok/s (-15%). Sometimes the fastest option when ROCm blocks.
- Arc A770 16 GB — Vulkan
- ~43 tok/s. Best value for money.
- Intel Xe iGPU (Lunar Lake)
- ~8–13 tok/s. Usable for a 2–3B model, but struggles with a 9B model.
#When to choose Vulkan
- You have an Intel Arc card
- Vulkan > oneAPI in 2026. Stable performance, no exotic installation.
- ROCm refuses to cooperate
- AMD card not officially supported, override crashes: Vulkan is the fallback that works.
- Multi-vendor setup
- A machine with GeForce + Intel Arc, a variable Thunderbolt+eGPU laptop. Vulkan tolerates everything.
- Tool distribution
- You're coding an app that must run on unknown machines: the Vulkan binary is the most compatible.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.