Intermediate 12 minGPU

Local LLM on an Intel Arc GPU (B580, A770) with Ollama and IPEX-LLM

Running a local LLM on an Intel Arc GPU is far from straightforward: the ecosystem is designed around CUDA. But with IPEX-LLM and its SYCL backend, the Intel Arc + Ollama combination becomes genuinely usable, and the VRAM-per-euro ratio of an A770 16 GB or a B580 12 GB is hard to beat. This guide shows how to install the stack, what you can actually expect in tokens/sec, and where the real limits are.

Choosing a machine? Our picks by budget →

By Thomas P.·Update 2026-08-27·Tested on Windows, macOS, and Linux
Recommended hardware

Good value for local AI: a GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395).

A mini PC is a complete machine: check the available memory and engine compatibility. It does not replace macOS/MLX or CUDA.

Why this choice? Our complete guide on GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395) →

Compare all options by budget, from €800 to €3,500 →

Small budget: RTX 5060 · Large models: RTX 5090 · Mac Studio.

On the go: which laptop for local AI →

Affiliate links — possible commission at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.

#Why Intel Arc for local AI?

In 2026, the expense that determines what you can run is VRAM. That is precisely where Intel Arc is aggressive: the Arc A770 has 16 GB of GDDR6 and is regularly available used for under €300, while the new Arc B580 (Battlemage architecture) offers 12 GB at a launch price of around €250. At this price, no new NVIDIA card offers as much memory.

The point is simple: VRAM determines the model size you can load without spilling onto the CPU. With 16 GB, a Qwen 3.5 9B or Granite 4.2 8B in Q4_K_M fits comfortably in memory, and a Gemma 4 12B in Q4 is perfectly usable. The tradeoff is software: Intel doesn't have CUDA, so you have to use its own oneAPI/SYCL stack and the IPEX-LLM library.

i
VRAM and model size
Quick reference for Q4_K_M: a 7B model uses ≈5 GB, a 14B uses ≈9 GB, and a 32B uses ≈19 GB. The 16 GB on an A770 comfortably covers up to 14B; the 12 GB on a B580 handles 7B–8B easily and runs 14B somewhat tight.

#B580 vs. A770: which one to choose

The two cards do not quite serve the same role. The B580 is newer, more energy-efficient, and better supported by current drivers, but it tops out at 12 GB. The A770 is older (Alchemist, 2022) and more power-hungry, but its 16 GB opens the door to larger models or a longer context.

Arc B580 (Battlemage)
12 GB GDDR6, ~€250 new. Better efficiency and more mature drivers. The right choice if you’re mainly targeting fast 7B–8B models.
Arc A770 (Alchemist)
16 GB GDDR6, often under €300 used. More VRAM for 14B models and extended context, at the cost of higher power consumption.
Arc A750 / B570
8 GB / 10 GB: troubleshooting is possible for a 7B model in tight Q4, but VRAM headroom is quickly exhausted. Best reserved for small budgets.
→
The deciding factor
If you’re hesitating, decide based on VRAM, not the name: 16 GB (A770) gives you more room to grow toward 14B and RAG than the 12 GB on a B580, even though the latter is newer.

#Hardware and driver requirements

Before installing anything, make sure you have a sound foundation. The most important point is Resizable BAR (ReBAR): on Arc GPUs, it is not optional. Without it, performance collapses and some loads fail.

  1. 01
    Enable Resizable BAR
    In the motherboard BIOS/UEFI, enable “Resizable BAR” (and “Above 4G Decoding”). This is essential for the Arc to work properly.
  2. 02
    Update the GPU drivers
    On Windows, install the latest Intel Arc drivers (Intel Arc Control). On Linux, use a recent kernel (6.2+) with the i915/xe driver and the Intel compute runtime (intel-opencl-icd, level-zero).
  3. 03
    Check detection
    Under Linux, the clinfo or sycl-ls command should list your Arc GPU as a Level Zero device. This proves that the oneAPI layer can see the card.
Terminal — check the GPU (Linux)
# Lister les périphériques SYCL visibles par oneAPI
sycl-ls

# Sortie attendue : une ligne [level_zero:gpu] ... Intel(R) Arc(TM) ...
# Si le GPU n'apparaît pas, le compute runtime n'est pas installé.
!
Resizable BAR required
This is the No. 1 cause of poor performance on Arc. If your tokens/sec are ridiculous or the model refuses to load on the GPU, check ReBAR in the BIOS before anything else.

#Install IPEX-LLM and the SYCL backend

IPEX-LLM is Intel’s library for optimizing LLM inference on its GPUs. It relies on oneAPI and llama.cpp’s SYCL backend, and provides a precompiled version of Ollama for Arc. This is the “IPEX-LLM” version of Ollama you need to use, not the standard Ollama binary, which only supports CUDA/ROCm/Metal.

The simplest route uses the ipex-llm[cpp] Python package, which installs an init-ollama command that generates Ollama binaries linked to the Intel backend. Create a dedicated environment so you don't break anything.

Terminal — install IPEX-LLM (conda/Linux)
# Environnement isolé
conda create -n ipex-llm python=3.11 -y
conda activate ipex-llm

# Installer IPEX-LLM avec le backend llama.cpp/Ollama
pip install --pre --upgrade ipex-llm[cpp]

# Générer les binaires Ollama optimisés Intel dans le dossier courant
mkdir ollama-arc && cd ollama-arc
init-ollama

On Windows, Intel also distributes a ready-to-use “Ollama portable” archive: extract it and run start-ollama.bat directly, without installing Python. It’s the fastest option for testing.

i
oneAPI already included
Recent ipex-llm[cpp] packages include the required oneAPI dependencies. You generally no longer need to install the oneAPI Base Toolkit separately or “source” setvars.sh for basic Ollama use.

#Run Ollama on Arc

Once the binaries are generated, launch the Ollama daemon as usual—it listens on http://localhost:11434—but set a few environment variables telling SYCL to use the GPU and offload everything to it.

Terminal — start the server
# Sélectionner le GPU Arc via Level Zero
export ONEAPI_DEVICE_SELECTOR=level_zero:0
# Cache de compilation SYCL (accélère les lancements suivants)
export SYCL_CACHE_PERSISTENT=1
# Forcer l'offload de toutes les couches sur le GPU
export OLLAMA_NUM_GPU=999

# Démarrer le daemon (depuis le dossier ollama-arc)
./ollama serve

In a second terminal, pull a model and start a conversation. Start small to validate the pipeline before loading a 12B model.

Terminal — first model
# Télécharger et lancer un modèle 8-9B en Q4_K_M
./ollama run qwen3.5:9b

# ou un modèle plus léger pour un premier test
./ollama run granite4.2:3b

On the first prompt, SYCL kernel compilation introduces a slight delay; thanks to SYCL_CACHE_PERSISTENT, subsequent launches are much faster. To confirm that the GPU is working, monitor its utilization with Intel's dedicated tool.

Terminal — monitor the GPU (Linux)
# Occupation et mémoire du GPU Intel en temps réel
sudo xpu-smi dump -d 0 -m 0,1,2

# Alternative légère si xpu-smi n'est pas installé
intel_gpu_top
→
Connect an interface
The daemon exposes the standard Ollama API on port 11434. You can therefore connect Open WebUI or LM Studio to it exactly as you would on a NVIDIA machine—nothing changes on the client side.

#Tokens/sec measured on B580 and A770

The figures below are approximate values observed on common models in Q4_K_M, with a short context and pure generation (excluding load time). They vary depending on the driver, the IPEX-LLM version, and cooling, but provide a realistic picture of what you can expect.

Granite 4.2 3B (Q4) — B580
≈ 45-55 tok/s. Smooth for chat and quick tasks.
Granite 4.2 8B (Q4) — B580
≈ 25–32 tok/s. Comfortable for everyday assistant use.
Qwen 3.5 9B (Q4) — A770
≈ 20-28 tok/s. One step below the B580 on the smaller format, but highly usable, with vision as a bonus.
Gemma 4 12B (Q4) — A770
≈ 11-16 tok/s. The 16 GB makes the difference: the multimodal model fits without CPU offload.
Gemma 4 12B (Q4) — B580
≈ 9–13 tok/s, near the 12 GB limit depending on the context; a long context can force offloading and cause throughput to drop.

The verdict is clear: with 8–9B models, both cards offer a pleasant real-time experience, often comparable to a RTX 3060 12 GB. At 12B, the A770 retains the advantage thanks to its VRAM. For context, the RTX 3060 12 GB remains the entry-level benchmark for NVIDIA; the Arc competes in this category, not against a RTX 4080.

i
Prompt processing
On Arc, prompt ingestion speed (prompt eval) is often the weak point, more than generation. For RAG with long contexts, expect a longer initial “thinking” time than on NVIDIA with the same VRAM.

#The limitations compared with NVIDIA

Let’s be honest: Arc is a smart budget option, not a no-compromise replacement for a CUDA card. Here’s what you need to know before buying.

Software ecosystem
You depend on IPEX-LLM and the SYCL backend. Some recent projects (new runtimes, llama.cpp features) arrive later than their CUDA counterparts. The “standard” version of Ollama isn't enough.
Supported quantizations
Standard GGUF formats work (Q4_K_M, Q5_K_M, Q8_0, FP16). However, some exotic or very recent quantizations may not be optimized and may even fall back to CPU computation.
Fine-tuning and advanced frameworks
Training/LoRA is possible through IPEX-LLM, but it is much less documented and equipped with fewer tools than NVIDIA. For serious fine-tuning, CUDA remains the preferred route.
Driver maturity
The situation is improving quickly, but a driver or IPEX-LLM version can introduce a regression. Pin a working version before updating blindly.
!
Who Arc is NOT the right choice for
If you want maximum compatibility—"it works on the first try"—intensive fine-tuning, or the ability to run the latest research projects, stick with NVIDIA. Arc rewards those willing to accept some tinkering in exchange for inexpensive VRAM.

#Common troubleshooting

The model runs on the CPU instead of the GPU
Check ONEAPI_DEVICE_SELECTOR=level_zero:0 and OLLAMA_NUM_GPU=999, and make sure sycl-ls lists the card. A disabled ReBAR can also force a fallback to the CPU.
Very low throughput despite an active GPU
Resizable BAR is almost always disabled in the BIOS. That's the first thing to check.
“out of memory” error with apparently sufficient VRAM
The context (num_ctx) increases memory usage. Reduce the context window, or switch to a lighter quantization (Q4 instead of Q5/Q8).
Very slow first prompt, then normal
SYCL kernel compilation. Make sure SYCL_CACHE_PERSISTENT=1 is exported so you don’t pay this cost again on every launch.
The GPU does not appear in sycl-ls
Missing compute runtime: install intel-opencl-icd and level-zero (Linux), or reinstall the Arc drivers (Windows).

#Go further

Once Ollama is operational on your Arc, the rest is the same as on any machine. These site guides build on this one:

Choose your quantization (Q4, Q5, Q8, FP16)
Critical on Arc, where VRAM is at a premium: understand the quality/memory trade-off to get the most from your 12 or 16 GB.
Choose your GPU for local AI
Put Arc in perspective against RTX cards and Macs to confirm that it’s the right purchase for your use case.
Install Ollama: Windows, macOS, and Linux
The basic installation guide and API usage on port 11434, useful for connecting Open WebUI behind your Arc.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.