Local LLM on an Intel Arc GPU (B580, A770) with Ollama and IPEX-LLM
Running a local LLM on an Intel Arc GPU is far from straightforward: the ecosystem is designed around CUDA. But with IPEX-LLM and its SYCL backend, the Intel Arc + Ollama combination becomes genuinely usable, and the VRAM-per-euro ratio of an A770 16 GB or a B580 12 GB is hard to beat. This guide shows how to install the stack, what you can actually expect in tokens/sec, and where the real limits are.
Choosing a machine? Our picks by budget →
Good value for local AI: a GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395).
A mini PC is a complete machine: check the available memory and engine compatibility. It does not replace macOS/MLX or CUDA.
Why this choice? Our complete guide on GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395) →
Compare all options by budget, from €800 to €3,500 →
Small budget: RTX 5060 · Large models: RTX 5090 · Mac Studio.
On the go: which laptop for local AI →
Affiliate links — possible commission at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.
#Why Intel Arc for local AI?
In 2026, the expense that determines what you can run is VRAM. That is precisely where Intel Arc is aggressive: the Arc A770 has 16 GB of GDDR6 and is regularly available used for under €300, while the new Arc B580 (Battlemage architecture) offers 12 GB at a launch price of around €250. At this price, no new NVIDIA card offers as much memory.
The point is simple: VRAM determines the model size you can load without spilling onto the CPU. With 16 GB, a Qwen 3.5 9B or Granite 4.2 8B in Q4_K_M fits comfortably in memory, and a Gemma 4 12B in Q4 is perfectly usable. The tradeoff is software: Intel doesn't have CUDA, so you have to use its own oneAPI/SYCL stack and the IPEX-LLM library.
#B580 vs. A770: which one to choose
The two cards do not quite serve the same role. The B580 is newer, more energy-efficient, and better supported by current drivers, but it tops out at 12 GB. The A770 is older (Alchemist, 2022) and more power-hungry, but its 16 GB opens the door to larger models or a longer context.
- Arc B580 (Battlemage)
- 12 GB GDDR6, ~€250 new. Better efficiency and more mature drivers. The right choice if you’re mainly targeting fast 7B–8B models.
- Arc A770 (Alchemist)
- 16 GB GDDR6, often under €300 used. More VRAM for 14B models and extended context, at the cost of higher power consumption.
- Arc A750 / B570
- 8 GB / 10 GB: troubleshooting is possible for a 7B model in tight Q4, but VRAM headroom is quickly exhausted. Best reserved for small budgets.
#Hardware and driver requirements
Before installing anything, make sure you have a sound foundation. The most important point is Resizable BAR (ReBAR): on Arc GPUs, it is not optional. Without it, performance collapses and some loads fail.
- 01Enable Resizable BARIn the motherboard BIOS/UEFI, enable “Resizable BAR” (and “Above 4G Decoding”). This is essential for the Arc to work properly.
- 02Update the GPU driversOn Windows, install the latest Intel Arc drivers (Intel Arc Control). On Linux, use a recent kernel (6.2+) with the i915/xe driver and the Intel compute runtime (intel-opencl-icd, level-zero).
- 03Check detectionUnder Linux, the clinfo or sycl-ls command should list your Arc GPU as a Level Zero device. This proves that the oneAPI layer can see the card.
#Install IPEX-LLM and the SYCL backend
IPEX-LLM is Intel’s library for optimizing LLM inference on its GPUs. It relies on oneAPI and llama.cpp’s SYCL backend, and provides a precompiled version of Ollama for Arc. This is the “IPEX-LLM” version of Ollama you need to use, not the standard Ollama binary, which only supports CUDA/ROCm/Metal.
The simplest route uses the ipex-llm[cpp] Python package, which installs an init-ollama command that generates Ollama binaries linked to the Intel backend. Create a dedicated environment so you don't break anything.
On Windows, Intel also distributes a ready-to-use “Ollama portable” archive: extract it and run start-ollama.bat directly, without installing Python. It’s the fastest option for testing.
#Run Ollama on Arc
Once the binaries are generated, launch the Ollama daemon as usual—it listens on http://localhost:11434—but set a few environment variables telling SYCL to use the GPU and offload everything to it.
In a second terminal, pull a model and start a conversation. Start small to validate the pipeline before loading a 12B model.
On the first prompt, SYCL kernel compilation introduces a slight delay; thanks to SYCL_CACHE_PERSISTENT, subsequent launches are much faster. To confirm that the GPU is working, monitor its utilization with Intel's dedicated tool.
#Tokens/sec measured on B580 and A770
The figures below are approximate values observed on common models in Q4_K_M, with a short context and pure generation (excluding load time). They vary depending on the driver, the IPEX-LLM version, and cooling, but provide a realistic picture of what you can expect.
- Granite 4.2 3B (Q4) — B580
- ≈ 45-55 tok/s. Smooth for chat and quick tasks.
- Granite 4.2 8B (Q4) — B580
- ≈ 25–32 tok/s. Comfortable for everyday assistant use.
- Qwen 3.5 9B (Q4) — A770
- ≈ 20-28 tok/s. One step below the B580 on the smaller format, but highly usable, with vision as a bonus.
- Gemma 4 12B (Q4) — A770
- ≈ 11-16 tok/s. The 16 GB makes the difference: the multimodal model fits without CPU offload.
- Gemma 4 12B (Q4) — B580
- ≈ 9–13 tok/s, near the 12 GB limit depending on the context; a long context can force offloading and cause throughput to drop.
The verdict is clear: with 8–9B models, both cards offer a pleasant real-time experience, often comparable to a RTX 3060 12 GB. At 12B, the A770 retains the advantage thanks to its VRAM. For context, the RTX 3060 12 GB remains the entry-level benchmark for NVIDIA; the Arc competes in this category, not against a RTX 4080.
#The limitations compared with NVIDIA
Let’s be honest: Arc is a smart budget option, not a no-compromise replacement for a CUDA card. Here’s what you need to know before buying.
- Software ecosystem
- You depend on IPEX-LLM and the SYCL backend. Some recent projects (new runtimes, llama.cpp features) arrive later than their CUDA counterparts. The “standard” version of Ollama isn't enough.
- Supported quantizations
- Standard GGUF formats work (Q4_K_M, Q5_K_M, Q8_0, FP16). However, some exotic or very recent quantizations may not be optimized and may even fall back to CPU computation.
- Fine-tuning and advanced frameworks
- Training/LoRA is possible through IPEX-LLM, but it is much less documented and equipped with fewer tools than NVIDIA. For serious fine-tuning, CUDA remains the preferred route.
- Driver maturity
- The situation is improving quickly, but a driver or IPEX-LLM version can introduce a regression. Pin a working version before updating blindly.
#Common troubleshooting
- The model runs on the CPU instead of the GPU
- Check ONEAPI_DEVICE_SELECTOR=level_zero:0 and OLLAMA_NUM_GPU=999, and make sure sycl-ls lists the card. A disabled ReBAR can also force a fallback to the CPU.
- Very low throughput despite an active GPU
- Resizable BAR is almost always disabled in the BIOS. That's the first thing to check.
- “out of memory” error with apparently sufficient VRAM
- The context (num_ctx) increases memory usage. Reduce the context window, or switch to a lighter quantization (Q4 instead of Q5/Q8).
- Very slow first prompt, then normal
- SYCL kernel compilation. Make sure SYCL_CACHE_PERSISTENT=1 is exported so you don’t pay this cost again on every launch.
- The GPU does not appear in sycl-ls
- Missing compute runtime: install intel-opencl-icd and level-zero (Linux), or reinstall the Arc drivers (Windows).
#Go further
Once Ollama is operational on your Arc, the rest is the same as on any machine. These site guides build on this one:
- Choose your quantization (Q4, Q5, Q8, FP16)
- Critical on Arc, where VRAM is at a premium: understand the quality/memory trade-off to get the most from your 12 or 16 GB.
- Choose your GPU for local AI
- Put Arc in perspective against RTX cards and Macs to confirm that it’s the right purchase for your use case.
- Install Ollama: Windows, macOS, and Linux
- The basic installation guide and API usage on port 11434, useful for connecting Open WebUI behind your Arc.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.