Which LLM on Radeon RX 9070 (16 GB) ?
The Radeon RX 9070 (RDNA 4, 16 GB GDDR6, 640 GB/s, 220 W) runs 20- to 24-billion-parameter models in Q4 and generates about 120 tokens per second on Llama 2 7B in Q4_0 under Vulkan, according to a public measurement. Under Linux, Ollama supports it with ROCm v7; under Windows, it is not on Ollama's ROCm list and uses Vulkan, which already delivers good results.
The RX 9070 is AMD's newest 16 GB card at this price point, with significantly lower power consumption than its predecessors in the RX 7000 series. This guide covers what it can run, what we know about its speed from public measurements, what changes between Linux and Windows, and where it stands against the RX 9070 XT and 16 GB RTX cards.
Choosing a machine? Our picks by budget →
For this setup: Radeon RX 9070 XT 16GB (ASUS Prime OC).
Why this choice? Our complete guide on Radeon RX 9070 XT 16GB (ASUS Prime OC) →
Compare all options by budget, from €800 to €3,500 →
On the go: which laptop for local AI →
Affiliate links — possible commission at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.
#What a RX 9070 does in 2026
The RX 9070 handles 8- to 14-billion-parameter models in Q4 without difficulty and accepts 20B to 24B models, with a short context for the largest ones. Its 16 GB of GDDR6 at 20.1 Gb/s on a 256-bit bus gives it 640 GB/s of bandwidth, correcting the previous version of this page (644). On Llama 2 7B in Q4_0, the public discussion in the llama.cpp repository reports 119.71 tokens per second for generation under Vulkan and 114.48 under ROCm, with prompt processing at 3,164 tokens per second under Vulkan, faster than the RX 7800 XT and 7900 GRE. Its other advantage is its 220 W TDP, 40 to 50 W lower than its older 16 GB counterparts.
- Architecture
- RDNA 4, a chip in the RX 9000 series, on the same 256-bit bus as the 9070 XT.
- Compute
- 3,584 stream processors (4,096 for the 9070 XT).
- Memory
- 16 GB GDDR6, 20.1 Gb/s, 256-bit bus, 640 GB/s.
- Power consumption
- 220 W TDP (304 W for the 9070 XT).
- Pricing
- Not specified: it changes every week; see /prix-gpu-ia.
#Linux, Windows: ROCm or Vulkan
This is the distinctive feature of RDNA 4 generation: support differs by operating system. AMD lists the RX 9070 in ROCm (gfx1201 target), and Ollama documentation cites it for Linux with the ROCm v7 driver. For Windows, however, Ollama's ROCm list stops at RX 7900 XTX; RX 9070 is not included. It remains usable through Vulkan, enabled by default, which also delivers good throughput. AMD's documentation finally specifies that Radeon RX 9000 are supported only on Ubuntu 24.04.4, Ubuntu 22.04.5, RHEL 10.1, and RHEL 9.7.
| Environment | Status | Practical consequence |
|---|---|---|
| Linux (Ubuntu 24.04.4, 22.04.5, RHEL 10.1, 9.7) | Official ROCm, gfx1201 target | Ollama with ROCm v7 driver |
| Windows | Absent from Ollama's ROCm list | Use Vulkan, enabled by default |
| Vulkan, all systems | General-purpose path | Comparable throughput to ROCm in public benchmarks |
#What fits in 16 GB of VRAM
Allow about 15 GB for the model and its KV cache when the card is not driving the display. The QuelLLM catalog lists the Q4 weights below; the context is added on top.
| Model | Weights | Headroom for context | Verdict |
|---|---|---|---|
| Llama 3.1 8B in Q4 | ≈ 6 GB | ≈ 9 GB | Very comfortable, very long context |
| Gemma 4 12B in Q8 | ≈ 13 GB | ≈ 2 GB | Fits, short context |
| gpt-oss 20B in Q4 | ≈ 13 GB | ≈ 2 GB | Holds up, short- to medium-length context |
| Devstral Small 2 24B in Q4 | ≈ 14 GB | ≈ 1 GB | Too cramped |
| Gemma 4 26B-A4B in Q4 | ≈ 16 GB | none | Too close to call |
These 16 GB are the boundary between the 24B, which fits, and the 26 to 30B models, which no longer fit. At this precise boundary, 16 GB cards from all brands are equivalent: the VRAM-by-page comparison covers them. What sets the RX 9070 apart is what it delivers beyond this capacity: faster prompt processing and lower power consumption.
#Measured throughput: Vulkan ahead of ROCm
The figures come from public discussions in the llama.cpp repository, all using Llama 2 7B in Q4_0 (3.56 GiB) and llama-bench; they are not measurements from the site. On the RX 9070, Vulkan delivers 3,164.10 tokens per second for prompt processing and 119.71 for generation; with Flash Attention, 2,859.98 and 119.51. ROCm delivers 2,381.77 and 114.48. The two backends are close for generation, with Vulkan slightly ahead, and the gap is more pronounced for prompt processing. Configurations, drivers, and commits differ from one series to another: treat these gaps as orders of magnitude.
| Backend | Flash Attention | pp512 processing | Generation tg128 |
|---|---|---|---|
| Vulkan | no | 3 164,10 t/s | 119,71 t/s |
| Vulkan | yes | 2 859,98 t/s | 119,51 t/s |
| ROCm | no | 2 381,77 t/s | 114,48 t/s |
The theoretical generation ceiling is 640 GB/s divided by 3.82 GB, or approximately 167 tokens per second; the card reaches 71% of that under Vulkan. Applying this throughput to larger models gives rough estimates. These are calculations, not measurements, and they assume that a larger model behaves like the reference 7B.
| Model (Q4) | Weights | Theoretical ceiling | Order of magnitude |
|---|---|---|---|
| Llama 3.1 8B | ≈ 6 GB | ≈ 107 tokens/s | ≈ 76 tokens/s |
| Gemma 4 12B | ≈ 7 GB | ≈ 91 tokens/s | ≈ 65 tokens/s |
| Phi-4 14B | ≈ 9 GB | ≈ 71 tokens/s | ≈ 50 tokens/s |
| Devstral Small 2 24B | ≈ 14 GB | ≈ 46 tokens/s | ≈ 32 tokens/s |
#What RDNA 4 changes for a local LLM
RDNA 4 adds dedicated AI hardware: the RX 9070 includes 112 AI accelerators according to the series spec sheet, versus 128 for the 9070 XT. You might expect much higher generation throughput than RDNA 3. Public benchmarks do not show that: 119.71 tokens per second for the RX 9070 under Vulkan, versus 118.27 for the RX 7800 XT in the same discussion. LLM generation is limited by memory bandwidth, and 640 GB/s is not far from the 624 GB/s of the 7800 XT. AI accelerators help more with prompt processing, where the RX 9070 is 57% faster than the 7800 XT.
The consequence for buyers: don't pay for RDNA 4 for conversational speed; pay for it to process prompts, consume less power, and extend software support, which will last longer for newer cards than for older ones. If your use case is chatting with an 8B to 14B model, a used RX 7800 XT delivers similar generation results.
#Compared with RX 9070 XT: the gap is larger than “10%”
The previous version of this page claimed that a RX 9070 was 10 to 12% slower than the 9070 XT. On the reference 7B model under Vulkan, the 9070 XT generates at 137.11 tokens per second versus 119.71, or 15% faster, and processes prompts at 5,036 tokens per second versus 3,164, or 59% faster. Yet both cards have the same memory, 16 GB, and the same bandwidth according to their specifications. The difference therefore comes from compute (4,096 stream processors versus 3,584), which has little impact on generation and a major impact on prompt processing.
| Criterion | RX 9070 | RX 9070 XT |
|---|---|---|
| Stream processors | 3 584 | 4 096 |
| TDP | 220 W | 304 W |
| pp512 processing | 3 164,10 t/s | 5 036,04 t/s |
| Generation tg128 | 119,71 t/s | 137,11 t/s |
In plain terms: for chat, the 9070 loses 15% speed in exchange for 84 W less power, a reasonable tradeoff. For RAG and long documents, the 9070 XT reads about 1.6 times faster.
#Against 16 GB RTX cards: the prompt remains the weak point
Compared with the NVIDIA 16 GB cards measured under Vulkan, the RX 9070 generates slightly more slowly, 8% below the RTX 4070 Ti Super and 12% below the RTX 5070 Ti, and reads the prompt at roughly half the speed. The RTX 5070 12 GB, often cited as a direct competitor, does not appear in the reference Vulkan discussion: there is therefore no reliable numerical comparison with it, only the capacity difference, 16 GB versus 12 GB.
| Card | pp512 processing | Generation tg128 |
|---|---|---|
| RX 9070 | 3 164,10 t/s | 119,71 t/s |
| RTX 4070 Ti Super | 6 099,18 t/s | 129,45 t/s |
| RTX 5070 Ti | 6 213,63 t/s | 135,63 t/s |
#Getting started
- 01Choose the systemOn Linux, use Ubuntu 24.04.4, Ubuntu 22.04.5, RHEL 10.1, or RHEL 9.7 to stay within AMD's support scope. On Windows, rely on Vulkan.
- 02Install OllamaOn Linux, install the ROCm v7 driver with amdgpu-install, then Ollama. On Windows, install Ollama: Vulkan is enabled by default.
- 03Load a modelChoose a 12B or 20B before a 24B, then verify with ollama ps that the model is running 100% on the GPU.
- 04Compare the backendsRun llama-bench with Vulkan and then ROCm on your model: in public benchmarks, Vulkan is slightly ahead.
Should you wait for a next generation? Nothing in the cited measurements justifies waiting for 8B to 24B use cases: the limit is 16 GB of capacity, which only a 20 to 24 GB card removes. If you already know you want 30B-and-larger models, skip this card; otherwise, buying a 16 GB card is justified by what you run today, not by what will be released tomorrow.
#2026 verdict
- A good choice
- If you want 16 GB in a modest 220 W card and accept Vulkan on Windows.
- Prefer the 9070 XT
- If your prompts are long: reading is 59% faster, for 84 W more.
- Prefer a 16 GB RTX
- If you need CUDA or twice-as-fast prompt processing.
Prices change quickly: the site tracker records the lowest price of graphics cards for local AI every Monday and Thursday, along with the price per GB of VRAM. This last figure should guide your comparison with the RTX 5070 or the RX 9070 XT.
- Graphics card prices for local AI, updated twice a week
- Ollama documentation: supported AMD cards
- ROCm GPU compatibility, AMD documentation
- Vulkan scoreboard for the llama.cpp repository
#Frequently asked questions
RX 9070 or RX 9070 XT for a local LLM?+
RX 9070 or RTX 5070 for a local LLM?+
Is RX 9070 supported by Ollama?+
How many tokens per second does a RX 9070 produce?+
Which models can you run on a RX 9070 with 16 GB?+
Vulkan or ROCm on RX 9070?+
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.