Intermediate 11 minRadeon RX 9000

Which LLM on Radeon RX 9070 (16 GB) ?

Direct response

The Radeon RX 9070 (RDNA 4, 16 GB GDDR6, 640 GB/s, 220 W) runs 20- to 24-billion-parameter models in Q4 and generates about 120 tokens per second on Llama 2 7B in Q4_0 under Vulkan, according to a public measurement. Under Linux, Ollama supports it with ROCm v7; under Windows, it is not on Ollama's ROCm list and uses Vulkan, which already delivers good results.

The RX 9070 is AMD's newest 16 GB card at this price point, with significantly lower power consumption than its predecessors in the RX 7000 series. This guide covers what it can run, what we know about its speed from public measurements, what changes between Linux and Windows, and where it stands against the RX 9070 XT and 16 GB RTX cards.

Choosing a machine? Our picks by budget →

By Mohamed Meguedmi·Update 2026-09-30·Tested on Windows, macOS, and Linux
Recommended hardware

For this setup: Radeon RX 9070 XT 16GB (ASUS Prime OC).

Why this choice? Our complete guide on Radeon RX 9070 XT 16GB (ASUS Prime OC) →

Compare all options by budget, from €800 to €3,500 →

On the go: which laptop for local AI →

Affiliate links — possible commission at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.

#What a RX 9070 does in 2026

The RX 9070 handles 8- to 14-billion-parameter models in Q4 without difficulty and accepts 20B to 24B models, with a short context for the largest ones. Its 16 GB of GDDR6 at 20.1 Gb/s on a 256-bit bus gives it 640 GB/s of bandwidth, correcting the previous version of this page (644). On Llama 2 7B in Q4_0, the public discussion in the llama.cpp repository reports 119.71 tokens per second for generation under Vulkan and 114.48 under ROCm, with prompt processing at 3,164 tokens per second under Vulkan, faster than the RX 7800 XT and 7900 GRE. Its other advantage is its 220 W TDP, 40 to 50 W lower than its older 16 GB counterparts.

Architecture
RDNA 4, a chip in the RX 9000 series, on the same 256-bit bus as the 9070 XT.
Compute
3,584 stream processors (4,096 for the 9070 XT).
Memory
16 GB GDDR6, 20.1 Gb/s, 256-bit bus, 640 GB/s.
Power consumption
220 W TDP (304 W for the 9070 XT).
Pricing
Not specified: it changes every week; see /prix-gpu-ia.

#Linux, Windows: ROCm or Vulkan

This is the distinctive feature of RDNA 4 generation: support differs by operating system. AMD lists the RX 9070 in ROCm (gfx1201 target), and Ollama documentation cites it for Linux with the ROCm v7 driver. For Windows, however, Ollama's ROCm list stops at RX 7900 XTX; RX 9070 is not included. It remains usable through Vulkan, enabled by default, which also delivers good throughput. AMD's documentation finally specifies that Radeon RX 9000 are supported only on Ubuntu 24.04.4, Ubuntu 22.04.5, RHEL 10.1, and RHEL 9.7.

Support for RX 9070
EnvironmentStatusPractical consequence
Linux (Ubuntu 24.04.4, 22.04.5, RHEL 10.1, 9.7)Official ROCm, gfx1201 targetOllama with ROCm v7 driver
WindowsAbsent from Ollama's ROCm listUse Vulkan, enabled by default
Vulkan, all systemsGeneral-purpose pathComparable throughput to ROCm in public benchmarks
i
One difference to know about
The target table in the Ollama documentation lists gfx1200 for the RX 9070 and gfx1201 for the 9070 XT as an example, while AMD's table associates the RX 9070 with gfx1201. Trust the rocminfo command on your machine, not a documentation example.

#What fits in 16 GB of VRAM

Allow about 15 GB for the model and its KV cache when the card is not driving the display. The QuelLLM catalog lists the Q4 weights below; the context is added on top.

What a RX 9070 accepts (weights only)
ModelWeightsHeadroom for contextVerdict
Llama 3.1 8B in Q4≈ 6 GB≈ 9 GBVery comfortable, very long context
Gemma 4 12B in Q8≈ 13 GB≈ 2 GBFits, short context
gpt-oss 20B in Q4≈ 13 GB≈ 2 GBHolds up, short- to medium-length context
Devstral Small 2 24B in Q4≈ 14 GB≈ 1 GBToo cramped
Gemma 4 26B-A4B in Q4≈ 16 GBnoneToo close to call

These 16 GB are the boundary between the 24B, which fits, and the 26 to 30B models, which no longer fit. At this precise boundary, 16 GB cards from all brands are equivalent: the VRAM-by-page comparison covers them. What sets the RX 9070 apart is what it delivers beyond this capacity: faster prompt processing and lower power consumption.

#Measured throughput: Vulkan ahead of ROCm

The figures come from public discussions in the llama.cpp repository, all using Llama 2 7B in Q4_0 (3.56 GiB) and llama-bench; they are not measurements from the site. On the RX 9070, Vulkan delivers 3,164.10 tokens per second for prompt processing and 119.71 for generation; with Flash Attention, 2,859.98 and 119.51. ROCm delivers 2,381.77 and 114.48. The two backends are close for generation, with Vulkan slightly ahead, and the gap is more pronounced for prompt processing. Configurations, drivers, and commits differ from one series to another: treat these gaps as orders of magnitude.

RX 9070 on Llama 2 7B Q4_0 (public llama.cpp measurements)
BackendFlash Attentionpp512 processingGeneration tg128
Vulkanno3 164,10 t/s119,71 t/s
Vulkanyes2 859,98 t/s119,51 t/s
ROCmno2 381,77 t/s114,48 t/s

The theoretical generation ceiling is 640 GB/s divided by 3.82 GB, or approximately 167 tokens per second; the card reaches 71% of that under Vulkan. Applying this throughput to larger models gives rough estimates. These are calculations, not measurements, and they assume that a larger model behaves like the reference 7B.

Estimated at 71% of the 640 GB/s ceiling (calculation, not measurement)
Model (Q4)WeightsTheoretical ceilingOrder of magnitude
Llama 3.1 8B≈ 6 GB≈ 107 tokens/s≈ 76 tokens/s
Gemma 4 12B≈ 7 GB≈ 91 tokens/s≈ 65 tokens/s
Phi-4 14B≈ 9 GB≈ 71 tokens/s≈ 50 tokens/s
Devstral Small 2 24B≈ 14 GB≈ 46 tokens/s≈ 32 tokens/s

#What RDNA 4 changes for a local LLM

RDNA 4 adds dedicated AI hardware: the RX 9070 includes 112 AI accelerators according to the series spec sheet, versus 128 for the 9070 XT. You might expect much higher generation throughput than RDNA 3. Public benchmarks do not show that: 119.71 tokens per second for the RX 9070 under Vulkan, versus 118.27 for the RX 7800 XT in the same discussion. LLM generation is limited by memory bandwidth, and 640 GB/s is not far from the 624 GB/s of the 7800 XT. AI accelerators help more with prompt processing, where the RX 9070 is 57% faster than the 7800 XT.

The consequence for buyers: don't pay for RDNA 4 for conversational speed; pay for it to process prompts, consume less power, and extend software support, which will last longer for newer cards than for older ones. If your use case is chatting with an 8B to 14B model, a used RX 7800 XT delivers similar generation results.

#Compared with RX 9070 XT: the gap is larger than “10%”

The previous version of this page claimed that a RX 9070 was 10 to 12% slower than the 9070 XT. On the reference 7B model under Vulkan, the 9070 XT generates at 137.11 tokens per second versus 119.71, or 15% faster, and processes prompts at 5,036 tokens per second versus 3,164, or 59% faster. Yet both cards have the same memory, 16 GB, and the same bandwidth according to their specifications. The difference therefore comes from compute (4,096 stream processors versus 3,584), which has little impact on generation and a major impact on prompt processing.

RX 9070 and RX 9070 XT under Vulkan (Llama 2 7B Q4_0, without FA)
CriterionRX 9070RX 9070 XT
Stream processors3 5844 096
TDP220 W304 W
pp512 processing3 164,10 t/s5 036,04 t/s
Generation tg128119,71 t/s137,11 t/s

In plain terms: for chat, the 9070 loses 15% speed in exchange for 84 W less power, a reasonable tradeoff. For RAG and long documents, the 9070 XT reads about 1.6 times faster.

#Against 16 GB RTX cards: the prompt remains the weak point

Compared with the NVIDIA 16 GB cards measured under Vulkan, the RX 9070 generates slightly more slowly, 8% below the RTX 4070 Ti Super and 12% below the RTX 5070 Ti, and reads the prompt at roughly half the speed. The RTX 5070 12 GB, often cited as a direct competitor, does not appear in the reference Vulkan discussion: there is therefore no reliable numerical comparison with it, only the capacity difference, 16 GB versus 12 GB.

RX 9070 and 16 GB RTX under Vulkan (Llama 2 7B Q4_0, without FA)
Cardpp512 processingGeneration tg128
RX 90703 164,10 t/s119,71 t/s
RTX 4070 Ti Super6 099,18 t/s129,45 t/s
RTX 5070 Ti6 213,63 t/s135,63 t/s

#Getting started

  1. 01
    Choose the system
    On Linux, use Ubuntu 24.04.4, Ubuntu 22.04.5, RHEL 10.1, or RHEL 9.7 to stay within AMD's support scope. On Windows, rely on Vulkan.
  2. 02
    Install Ollama
    On Linux, install the ROCm v7 driver with amdgpu-install, then Ollama. On Windows, install Ollama: Vulkan is enabled by default.
  3. 03
    Load a model
    Choose a 12B or 20B before a 24B, then verify with ollama ps that the model is running 100% on the GPU.
  4. 04
    Compare the backends
    Run llama-bench with Vulkan and then ROCm on your model: in public benchmarks, Vulkan is slightly ahead.
Terminal
ollama ps
rocminfo | grep -i gfx
./llama-bench -m llama-2-7b.Q4_0.gguf -ngl 100 -fa 0,1

Should you wait for a next generation? Nothing in the cited measurements justifies waiting for 8B to 24B use cases: the limit is 16 GB of capacity, which only a 20 to 24 GB card removes. If you already know you want 30B-and-larger models, skip this card; otherwise, buying a 16 GB card is justified by what you run today, not by what will be released tomorrow.

#2026 verdict

A good choice
If you want 16 GB in a modest 220 W card and accept Vulkan on Windows.
Prefer the 9070 XT
If your prompts are long: reading is 59% faster, for 84 W more.
Prefer a 16 GB RTX
If you need CUDA or twice-as-fast prompt processing.

Prices change quickly: the site tracker records the lowest price of graphics cards for local AI every Monday and Thursday, along with the price per GB of VRAM. This last figure should guide your comparison with the RTX 5070 or the RX 9070 XT.

#Frequently asked questions

FAQ
RX 9070 or RX 9070 XT for a local LLM?+
The 9070 XT generates 15% faster and reads prompts 59% faster on the reference 7B model under Vulkan, at the cost of a 304 W TDP versus 220 W. For chat, the 9070 is sufficient; for RAG and long documents, the XT is justified. Compare current prices.
RX 9070 or RTX 5070 for a local LLM?+
The RX 9070 provides 16 GB versus 12 GB, opening the door to 20- to 24-billion-parameter models. The RTX 5070 retains the CUDA ecosystem and offers faster prompt processing. No public Vulkan benchmark for the RTX 5070 was found, so a numerical comparison remains to be done.
Is RX 9070 supported by Ollama?+
On Linux, yes, with the ROCm v7 driver: it appears in the Ollama list. On Windows, it is not in the Ollama ROCm list, which stops at the RX 7900 XTX, but Vulkan, enabled by default, lets you use it. Check with ollama ps that the model is loaded on the GPU.
How many tokens per second does a RX 9070 produce?+
For Llama 2 7B in Q4_0, public discussions in the llama.cpp repository report 119.71 tokens per second under Vulkan and 114.48 under ROCm during generation. For a 24B model in Q4, the calculation gives about 32 tokens per second: an estimate to measure on your system.
Which models can you run on a RX 9070 with 16 GB?+
Llama 3.1 8B, Gemma 4 12B up to Q8, gpt-oss 20B and Devstral Small 2 24B in Q4, with the latter two using a short context. A 26B to 30B model in Q4 exceeds 16 GB of VRAM and spills onto the processor, with a significant drop in speed.
Vulkan or ROCm on RX 9070?+
In the public llama.cpp benchmarks, Vulkan generates slightly faster (119.71 versus 114.48 tokens per second) and reads the prompt faster (3,164 versus 2,382). On Windows, Vulkan is the default path anyway. Test both with your model before deciding.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.