Ollama + AMD GPU (ROCm): configure Radeon RX 6000 to 9000
AMD lagged behind CUDA for a long time, but ROCm 6 has caught up enough for Radeon RX 6000, 7000, and 9000 to be fully usable for LLM inference. This guide covers a clean installation, Ollama configuration, and the override that lets you run cards not officially listed.
#Why ROCm
To use an AMD GPU, Ollama (which relies on llama.cpp) must run through ROCm — AMD’s equivalent of CUDA. Without ROCm, the card is invisible to inference and everything runs on the CPU.
#Officially compatible cards
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
- RX 7900 XTX / XT / GRE
- gfx1100. Officially supported. About 80% of the performance of a RTX 4080.
- RX 7800 / 7700 XT
- gfx1101. Official support since ROCm 6.1.
- RX 6900 / 6800 / 6700 XT
- gfx1030/1031. Official support for large cards, override for smaller ones.
- Instinct MI200 / MI300
- Data center cards. Full native support.
#1. Install ROCm
The two best-supported distributions are Ubuntu 22.04/24.04 and RHEL/Fedora. Debian, Arch, and similar distributions require additional workarounds.
#2. Ollama and ROCm
Ollama 0.3+ detects ROCm automatically. If you installed Ollama BEFORE ROCm, reinstall it—the script will then detect the AMD GPU and install the correct libraries.
#3. Force an unsupported card
If rocminfo does not list your GPU, or if Ollama continues to use the CPU, the HSA_OVERRIDE_GFX_VERSION variable forces ROCm to pretend.
- RX 6600 / 6700
- HSA_OVERRIDE_GFX_VERSION=10.3.0
- RX 5700 XT
- HSA_OVERRIDE_GFX_VERSION=10.1.0 (fragile support)
- RX 6800 / 6900
- Supported natively; no override needed.
#4. Expected performance
For Qwen 3.5 9B Q4_K_M (6.6 GB, the reference 8 GB choice in 2026), here are the approximate tokens-per-second figures during inference:
- RX 7900 XTX (24 GB)
- ~90 tok/s. Comparable to a RTX 4080.
- RX 7800 XT (16 GB)
- ~62 tok/s. Between RTX 4070 and 4070 Ti.
- RX 6900 XT (16 GB)
- ~52 tok/s. Equivalent to RTX 3080 10 GB.
- RX 6700 XT (12 GB, override)
- ~33 tok/s. Between RTX 3060 and 3060 Ti.
#Troubleshooting
- rocminfo is empty
- Drivers not installed or user not in the render group. Check groups $USER.
- Ollama uses the CPU despite ROCm
- Often an architecture incompatibility. Try the HSA_OVERRIDE corresponding to the generation above.
- Crash on large model (OOM)
- Use rocm-smi to see how much VRAM is being used. If you exceed the limit, reduce num_ctx or switch to Q3_K_M.
- Incompatible ROCm version Ollama
- Ollama packaged for ROCm 6.x. ROCm 5 does not work. Update.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.