Troubleshoot Ollama: GPU not detected, slowdowns, errors memory
You installed Ollama, you have a good GPU, and yet responses are crawling—a classic sign that an undetected ollama gpu has shifted the model to the CPU. This guide reviews the most common failures (CPU fallback, out of memory, slowdowns, driver conflicts) and gives you diagnostic commands each time to find the real cause instead of guessing.
#Typical symptoms
Ollama rarely fails outright: most often it “works,” but poorly. The daemon is listening on http://localhost:11434, the model responds, but something is off. Knowing how to recognize the symptom already points you toward the right family of causes.
- Very slow responses
- A few tokens per second when you expect dozens: the model is probably running on the CPU, with the GPU undetected or unused.
- GPU at 0% utilization
- nvidia-smi or rocm-smi show an idle card during generation: Ollama did not use it.
- Out-of-memory error
- Loading fails or the model is evicted with a CUDA/HIP “out of memory” message: the model or its context exceeds VRAM.
- Sudden slowdown after an update
- A throughput collapse from one day to the next points to a driver, an Ollama version, or a CUDA/ROCm conflict.
- Partial offload
- Some layers on the GPU, the rest on the CPU: it works, but is much slower than a full offload.
#Verify that the GPU is actually being used
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
Before looking for a solution, establish one fact: does Ollama use the GPU, yes or no? The most direct command is ollama ps, which displays the loaded models and, above all, the CPU/GPU split while a model is in memory.
The PROCESSOR column is the key. “100% GPU” means the entire model is on the card—that is what you want. “100% CPU” confirms a full fallback. “60%/40% CPU/GPU” indicates partial offloading: the model does not fit entirely in VRAM.
To cross-check the information, monitor the GPU on the manufacturer's side during generation. If usage rises, the GPU is working; if it stays at zero, Ollama ignores it.
#Read Ollama logs
The server logs state exactly which backend was loaded and why a GPU was selected or rejected. They are the source of truth when ollama ps and nvidia-smi contradict each other.
Look for lines mentioning “inference compute,” “library=cuda,” or “library=rocm,” along with the number of offloaded layers (“offloaded X/Y layers to GPU”). If you see “no compatible GPUs were discovered” or “library=cpu,” you’ve found the cause: Ollama detected no usable GPU.
#CPU fallback and its common causes
When Ollama can't use the GPU, it doesn't crash: it silently switches to the CPU. This is the most confusing behavior because “everything works,” apparently. Here are the most common causes, from the most mundane to the most subtle.
- 01Not enough VRAM for the modelIf the model (weights + context) exceeds the available VRAM, Ollama offloads some of it to the CPU, or even all of it. A 14B model in Q4 requires ≈9 GB, while a 32B model requires ≈19 GB: on a 12 GB card, the 32B model will never fit entirely.
- 02GPU driver missing or too oldWithout a recent NVIDIA driver (or ROCm on AMD), Ollama detects no compatible GPU and falls back to the CPU. This is the No. 1 cause of “ollama gpu not detected.”
- 03GPU occupied by another processAnother model, a game, a Python notebook, or another Ollama instance may already be using up the VRAM. nvidia-smi lists the processes and memory usage.
- 04Ollama in a container without GPU passthroughIn Docker, without --gpus all (NVIDIA) or the exposed ROCm devices, the container cannot see the card. The daemon runs, but on the CPU.
- 05Unsupported GPUCards that are too old (insufficient CUDA compute capability) or AMD GPUs outside the official ROCm list: Ollama deliberately ignores them.
#Out-of-memory errors: how to read and fix them
A “CUDA error: out of memory” message (or “HIP out of memory” on AMD) means the model requires more VRAM than is available. Two factors increase this requirement: model size and context window. Reducing either one resolves most cases.
- Model too large
- Switch to a lighter quantization (Q4_K_M instead of Q5/Q8) or a smaller model. Q4 rule of thumb: 7B≈5 GB · 14B≈9 GB · 32B≈19 GB · 70B≈40 GB.
- Context too long
- A large num_ctx multiplies KV cache memory usage. Reducing the context window (num_ctx) frees up a lot of VRAM, especially on large models.
- VRAM fragmented by another process
- Close games, notebooks, and other models. Restart the Ollama daemon to start with clean VRAM.
- Resident memory
- A previously loaded model still occupies VRAM. ollama stop <modèle> unloads it immediately instead of waiting for keep_alive.
#NVIDIA/AMD drivers up to date
Many “ollama gpu not detected” problems can be fixed by upgrading the driver layer. Ollama bundles its own CUDA/ROCm libraries, but it depends on the system driver to communicate with the hardware.
On the AMD side, ROCm installation is more demanding: the card must appear in the list of supported GPUs, and you may sometimes need to export a variable to force support for a nearby architecture (HSA_OVERRIDE_GFX_VERSION). First, verify that rocminfo can see the GPU.
#Sudden slowdowns
A throughput rate that was good and then suddenly drops almost always has an identifiable cause. Unlike permanent CPU fallback, sudden slowness indicates a recent state change.
- Newly added partial offload
- You increased the context or loaded a larger model: part of it is being offloaded to the CPU. Check ollama ps and the “offloaded N/M layers” line.
- Thermal throttling
- A GPU that overheats reduces its clock speed. nvidia-smi displays the temperature and “P-State” status; above ~83 °C, the card throttles.
- Update of Ollama or the driver
- A regression can slow inference. Note the working version (ollama --version) before updating so you can roll back.
- Shared VRAM / system memory
- On Windows, when VRAM overflows, the driver can use system RAM (shared GPU memory), which causes performance to collapse. Reduce the model instead of letting it overflow.
- Reloading on every request
- A keep_alive value that is too short reloads the model constantly. Increase OLLAMA_KEEP_ALIVE if you chain requests.
#Diagnostic checklist
When something seems wrong, go through these steps in order: they run from the most common to the rarest and keep you from heading in the wrong direction.
- 011. Confirm the symptomollama ps pendant une génération : la colonne PROCESSOR indique-t-elle CPU, GPU ou un mélange ?
- 022. Verify that the GPU exists for the systemDoes nvidia-smi (or rocm-smi) respond? If not, the problem is with the driver, not Ollama.
- 033. Read the logsjournalctl -u ollama -f (or OLLAMA_DEBUG=1 ollama serve): look for “library=cuda/rocm/cpu” and “offloaded N/M layers”.
- 044. Check available VRAMnvidia-smi: is another process using the card? Does the model fit in the available VRAM?
- 055. Reduce the load if OOM or partially offloadedLighter quantization, a smaller model, or reduced num_ctx. Unload unused models with ollama stop.
- 066. Update and restartKeep the driver and Ollama up to date, then restart the service (and the machine after a driver change).
#Go further
Once the GPU is detected correctly, these site guides help you get the most from your setup and prevent the problems from coming back:
- Choose your quantization (Q4, Q5, Q8, FP16)
- The number 1 lever against memory errors: understand the quality/VRAM tradeoff so you can choose a format that fits on your card.
- Run an LLM locally without a GPU (CPU only)
- If your hardware does not support GPU offload, this guide shows how to remain usable on pure CPU and which models to choose based on RAM.
- Install Ollama: Windows, macOS, and Linux
- Review the clean installation of the daemon and drivers, often the real source of an undetected GPU.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.