Intermediate 11 minOllama

Troubleshoot Ollama: GPU not detected, slowdowns, errors memory

You installed Ollama, you have a good GPU, and yet responses are crawling—a classic sign that an undetected ollama gpu has shifted the model to the CPU. This guide reviews the most common failures (CPU fallback, out of memory, slowdowns, driver conflicts) and gives you diagnostic commands each time to find the real cause instead of guessing.

By Marie L.·Update 2026-08-27·Tested on Windows, macOS, and Linux

#Typical symptoms

Ollama rarely fails outright: most often it “works,” but poorly. The daemon is listening on http://localhost:11434, the model responds, but something is off. Knowing how to recognize the symptom already points you toward the right family of causes.

Very slow responses
A few tokens per second when you expect dozens: the model is probably running on the CPU, with the GPU undetected or unused.
GPU at 0% utilization
nvidia-smi or rocm-smi show an idle card during generation: Ollama did not use it.
Out-of-memory error
Loading fails or the model is evicted with a CUDA/HIP “out of memory” message: the model or its context exceeds VRAM.
Sudden slowdown after an update
A throughput collapse from one day to the next points to a driver, an Ollama version, or a CUDA/ROCm conflict.
Partial offload
Some layers on the GPU, the rest on the CPU: it works, but is much slower than a full offload.

#Verify that the GPU is actually being used

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

Before looking for a solution, establish one fact: does Ollama use the GPU, yes or no? The most direct command is ollama ps, which displays the loaded models and, above all, the CPU/GPU split while a model is in memory.

Terminal — loaded model status
# Charger un modèle puis, dans un autre terminal, l'interroger
ollama run qwen3.5:9b "bonjour"

# Pendant que le modèle est en mémoire, vérifier la répartition
ollama ps

The PROCESSOR column is the key. “100% GPU” means the entire model is on the card—that is what you want. “100% CPU” confirms a full fallback. “60%/40% CPU/GPU” indicates partial offloading: the model does not fit entirely in VRAM.

i
Read the PROCESSOR column
ollama ps n'affiche la répartition que tant qu'un modèle est chargé en mémoire. Par défaut Ollama décharge un modèle après 5 minutes d'inactivité (keep_alive) — lancez la commande juste après une génération.

To cross-check the information, monitor the GPU on the manufacturer's side during generation. If usage rises, the GPU is working; if it stays at zero, Ollama ignores it.

Terminal — monitor the GPU
# NVIDIA : rafraîchissement toutes les secondes
nvidia-smi -l 1

# AMD (ROCm)
rocm-smi

# Vue plus lisible et interactive (si installé)
nvtop

#Read Ollama logs

The server logs state exactly which backend was loaded and why a GPU was selected or rejected. They are the source of truth when ollama ps and nvidia-smi contradict each other.

Terminal — view logs
# Linux (service systemd)
journalctl -u ollama -f

# macOS / lancement manuel : les logs vont sur la sortie du serveur
# Relancer le daemon en verbeux pour tout voir
OLLAMA_DEBUG=1 ollama serve

Look for lines mentioning “inference compute,” “library=cuda,” or “library=rocm,” along with the number of offloaded layers (“offloaded X/Y layers to GPU”). If you see “no compatible GPUs were discovered” or “library=cpu,” you’ve found the cause: Ollama detected no usable GPU.

→
The message that matters
The “offloaded N/M layers to GPU” line tells you everything: N=M is perfect, N<M is a partial offload (insufficient VRAM), and N=0 is CPU-only. Always start from this line when diagnosing.

#CPU fallback and its common causes

When Ollama can't use the GPU, it doesn't crash: it silently switches to the CPU. This is the most confusing behavior because “everything works,” apparently. Here are the most common causes, from the most mundane to the most subtle.

  1. 01
    Not enough VRAM for the model
    If the model (weights + context) exceeds the available VRAM, Ollama offloads some of it to the CPU, or even all of it. A 14B model in Q4 requires ≈9 GB, while a 32B model requires ≈19 GB: on a 12 GB card, the 32B model will never fit entirely.
  2. 02
    GPU driver missing or too old
    Without a recent NVIDIA driver (or ROCm on AMD), Ollama detects no compatible GPU and falls back to the CPU. This is the No. 1 cause of “ollama gpu not detected.”
  3. 03
    GPU occupied by another process
    Another model, a game, a Python notebook, or another Ollama instance may already be using up the VRAM. nvidia-smi lists the processes and memory usage.
  4. 04
    Ollama in a container without GPU passthrough
    In Docker, without --gpus all (NVIDIA) or the exposed ROCm devices, the container cannot see the card. The daemon runs, but on the CPU.
  5. 05
    Unsupported GPU
    Cards that are too old (insufficient CUDA compute capability) or AMD GPUs outside the official ROCm list: Ollama deliberately ignores them.
Terminal — what's using the VRAM?
# Voir les processus et la mémoire GPU consommée
nvidia-smi

# Repérer d'autres instances Ollama qui tourneraient déjà
ps aux | grep ollama
!
Docker and the GPU
A Ollama container running on the CPU even though the host machine has a GPU is almost always caused by missing passthrough. Check the --gpus all flag and the installation of the NVIDIA Container Toolkit on the host.

#Out-of-memory errors: how to read and fix them

A “CUDA error: out of memory” message (or “HIP out of memory” on AMD) means the model requires more VRAM than is available. Two factors increase this requirement: model size and context window. Reducing either one resolves most cases.

Model too large
Switch to a lighter quantization (Q4_K_M instead of Q5/Q8) or a smaller model. Q4 rule of thumb: 7B≈5 GB · 14B≈9 GB · 32B≈19 GB · 70B≈40 GB.
Context too long
A large num_ctx multiplies KV cache memory usage. Reducing the context window (num_ctx) frees up a lot of VRAM, especially on large models.
VRAM fragmented by another process
Close games, notebooks, and other models. Restart the Ollama daemon to start with clean VRAM.
Resident memory
A previously loaded model still occupies VRAM. ollama stop <modèle> unloads it immediately instead of waiting for keep_alive.
Terminal — reduce memory pressure
# Décharger un modèle tout de suite
ollama stop qwen3.6:35b

# Lancer avec un contexte réduit (via l'API ou un Modelfile)
# Exemple : forcer num_ctx à 4096 au lieu de la valeur par défaut
ollama run qwen3.5:9b
# puis dans le prompt interactif :
# /set parameter num_ctx 4096
→
Estimate before loading
Add the model size (based on its quantization) and some headroom for context. Aim to stay below ~90% of your total VRAM: beyond that, partial offloading or OOM may occur. On a RTX 4090 24 GB, a 32B Q4 (≈19 GB) fits; a 70B Q4 (≈40 GB) does not.

#NVIDIA/AMD drivers up to date

Many “ollama gpu not detected” problems can be fixed by upgrading the driver layer. Ollama bundles its own CUDA/ROCm libraries, but it depends on the system driver to communicate with the hardware.

Terminal — check the NVIDIA driver
# Doit afficher la version du driver et CUDA
nvidia-smi

# Si la commande échoue : le driver n'est pas installé ou pas chargé
# Vérifier que le module noyau est bien chargé (Linux)
lsmod | grep nvidia

On the AMD side, ROCm installation is more demanding: the card must appear in the list of supported GPUs, and you may sometimes need to export a variable to force support for a nearby architecture (HSA_OVERRIDE_GFX_VERSION). First, verify that rocminfo can see the GPU.

Terminal — check ROCm (AMD)
# Le GPU doit apparaître dans la liste des agents
rocminfo | grep -i gfx

# État et mémoire du GPU AMD
rocm-smi

# Forcer une architecture proche si la carte n'est pas officiellement listée
export HSA_OVERRIDE_GFX_VERSION=11.0.0
!
After a driver update, restart the service
A new driver takes effect only after the kernel module is reloaded—the safest approach is to restart the machine, then the Ollama service (sudo systemctl restart ollama). A daemon launched before the update continues to see the old state.

#Sudden slowdowns

A throughput rate that was good and then suddenly drops almost always has an identifiable cause. Unlike permanent CPU fallback, sudden slowness indicates a recent state change.

Newly added partial offload
You increased the context or loaded a larger model: part of it is being offloaded to the CPU. Check ollama ps and the “offloaded N/M layers” line.
Thermal throttling
A GPU that overheats reduces its clock speed. nvidia-smi displays the temperature and “P-State” status; above ~83 °C, the card throttles.
Update of Ollama or the driver
A regression can slow inference. Note the working version (ollama --version) before updating so you can roll back.
Shared VRAM / system memory
On Windows, when VRAM overflows, the driver can use system RAM (shared GPU memory), which causes performance to collapse. Reduce the model instead of letting it overflow.
Reloading on every request
A keep_alive value that is too short reloads the model constantly. Increase OLLAMA_KEEP_ALIVE if you chain requests.
Terminal — temperature and frequency
# Surveiller température, conso et fréquence en continu
nvidia-smi -l 1

# Garder les modèles chargés plus longtemps (ex : 30 min)
export OLLAMA_KEEP_ALIVE=30m
sudo systemctl restart ollama

#Diagnostic checklist

When something seems wrong, go through these steps in order: they run from the most common to the rarest and keep you from heading in the wrong direction.

  1. 01
    1. Confirm the symptom
    ollama ps pendant une génération : la colonne PROCESSOR indique-t-elle CPU, GPU ou un mélange ?
  2. 02
    2. Verify that the GPU exists for the system
    Does nvidia-smi (or rocm-smi) respond? If not, the problem is with the driver, not Ollama.
  3. 03
    3. Read the logs
    journalctl -u ollama -f (or OLLAMA_DEBUG=1 ollama serve): look for “library=cuda/rocm/cpu” and “offloaded N/M layers”.
  4. 04
    4. Check available VRAM
    nvidia-smi: is another process using the card? Does the model fit in the available VRAM?
  5. 05
    5. Reduce the load if OOM or partially offloaded
    Lighter quantization, a smaller model, or reduced num_ctx. Unload unused models with ollama stop.
  6. 06
    6. Update and restart
    Keep the driver and Ollama up to date, then restart the service (and the machine after a driver change).
i
Isolate before fixing
The golden rule: change only one thing at a time and recheck with ollama ps. Changing the driver, context, and model at the same time makes it impossible to identify what actually solved—or worsened—the problem.

#Go further

Once the GPU is detected correctly, these site guides help you get the most from your setup and prevent the problems from coming back:

Choose your quantization (Q4, Q5, Q8, FP16)
The number 1 lever against memory errors: understand the quality/VRAM tradeoff so you can choose a format that fits on your card.
Run an LLM locally without a GPU (CPU only)
If your hardware does not support GPU offload, this guide shows how to remain usable on pure CPU and which models to choose based on RAM.
Install Ollama: Windows, macOS, and Linux
Review the clean installation of the daemon and drivers, often the real source of an undetected GPU.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.