Which LLM on Radeon RX 9070 XT (16 GB) ?
The RX 9070 XT (16 GB GDDR6, up to 640 GB/s according to AMD) is a good card for local LLMs on Linux with ROCm 7, or through Vulkan on Windows. It loads models with up to 13–14 GB of Q4 weights without offloading, such as gpt-oss 20B or Mistral Small 24B. Dense models with 27 billion parameters are out of reach.
The RX 9070 XT is AMD's high-end RDNA 4 Radeon. This guide explains what its 16 GB enables, how to install it on Linux or Windows, what public benchmarks show, and when a RTX 5070 Ti or a 24 GB card is preferable.
Choosing a machine? Our picks by budget →
For this setup: Radeon RX 9070 XT 16GB (ASUS Prime OC).
Why this choice? Our complete guide on Radeon RX 9070 XT 16GB (ASUS Prime OC) →
Compare all options by budget, from €800 to €3,500 →
On the go: which laptop for local AI →
Affiliate links — possible commission at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.
#RX 9070 XT for LLMs: 16 GB, 640 GB/s, and strong compute
The Radeon RX 9070 XT is the most capable consumer Radeon for LLMs among the RX 9000: 16 GB of GDDR6 on a 256-bit bus, bandwidth of up to 640 GB/s according to AMD, and substantially greater matrix compute than the RX 7900 XTX in FP16 and FP8. It runs models whose weights remain below about 13 to 14 GB in Q4 without overflow: a 9B or 12B with a long context, gpt-oss 20B, Mistral Small 24B with a short context. Its limitation is capacity: dense models with 27 billion parameters in Q4 (about 16 GB) do not fit.
| Feature | Value |
|---|---|
| Memory | 16 GB GDDR6, 256-bit bus |
| Bandwidth | up to 640 GB/s |
| Stream processors / compute units | 4 096 / 64 |
| AI accelerators | 128 |
| FP16 / FP8 matrix multiplication | 195 / 389 TFLOPs |
| Typical card power | 304 W |
| Recommended minimum power supply | 750 W |
| Power connectors | 2 x 8-pin |
#Software: ROCm, Vulkan, LM Studio, vLLM
The RX 9070 XT appears in AMD's Radeon documentation matrix for ROCm 7.2.1 (Ubuntu 22.04, 24.04, and 25.10). Ollama lists it among the cards supported on Linux with the ROCm v7 driver and maps the gfx1201 target to the 9070 XT in its LLVM targets table. LM Studio announced support for AMD 9000-series GPUs on Linux with ROCm in version 0.3.19, released in July 2025. vLLM lists the Radeon RX 9000 (gfx1200 and gfx1201) among its AMD hardware.
Under Windows, the ROCm list from Ollama stops at the RX 7000: the RX 9000 are not included. Ollama nevertheless specifies that Vulkan provides additional support under Windows and Linux, enabled by default. A 9070 XT under Windows therefore effectively uses Vulkan; check with ollama ps that the model is actually on the GPU, and consult the llama.cpp Vulkan guide if you compile it yourself.
#Getting started on Linux with Ollama
- 01Install the ROCm driverFollow AMD’s ROCm documentation with the amdgpu-install utility, as Ollama requires, add your user to the render and video groups, then reboot.
- 02Confirm detectionrocminfo must list a GPU agent and rocm-smi must display 16 GB. Without that, Ollama will use the processor.
- 03Run your first modelInstall Ollama with its official script and start a model that fits in 16 GB, such as a 9B in Q4.
- 04Check placementAfter the first response, ollama ps should show 100% GPU. Shared GPU and processor usage indicates that the model or context is too large.
#Which models fit in 16 GB
The weights come from the QuelLLM catalog, in Q4 unless noted otherwise. Add the KV cache and about one gigabyte of headroom. The full list by tier is in the 16 GB VRAM guide.
| Model | Weights | Does it fit in 16 GB? | Theoretical ceiling |
|---|---|---|---|
| Qwen 3.5 9B Q4 | 6 GB | Yes, very broad | ≈ 107 tokens/s |
| Qwen 3.5 9B Q8 | 10 GB | Yes | 64 tokens/s |
| Gemma 4 12B Q4 | 7 GB | Yes, long context | ≈ 91 tokens/s |
| Gemma 4 12B Q8 | 13 GB | Yes, short context | ≈ 49 tokens/s |
| gpt-oss 20B (MoE) | 13 GB | Yes, moderate context | depends on the active weights |
| Mistral Small 24B Q4 | 14 GB | Tight | ≈ 46 tokens/s |
| Qwen 3.8 27B Q4 | 16 GB | No | - |
The ceiling is bandwidth divided by the weights read at each token; no card reaches it exactly. MoE models such as gpt-oss 20B read fewer weights per token than their total size and therefore run faster than a dense model of the same size, provided they fit entirely in VRAM.
#What the public measurements show
The llama.cpp community table on ROCm measures the same model, Llama 2 7B in Q4_0, on several cards. For the RX 9070 XT, one participant reports 5 055 tokens per second when reading the prompt (pp512) and 101 during generation (tg128), without Flash Attention. On the 7900 XTX, the table gives 3 552 and 167; on the 9060 XT, 1 420 and 68.
| Card | AMD bandwidth | Prompt processing | Generation |
|---|---|---|---|
| RX 7900 XTX | 960 GB/s | 3 552 t/s | 167 t/s |
| RX 9070 XT | 640 GB/s | 5 055 t/s | 101 t/s |
| RX 9060 XT | 320 GB/s | 1 420 t/s | 68 t/s |
Two takeaways. First, the 9070 XT reads a long prompt about 40% faster than the 7900 XTX, despite lower bandwidth: that's the advantage of more powerful matrix computation, which is useful for RAG, long documents, and coding agents that return a lot of context. Second, during generation, the 7900 XTX remains ahead, as its 50% higher bandwidth would suggest.
#Which model to choose for each use case
With 16 GB, the right choice depends less on the maximum size than on the headroom left for context. A simple rule: keep at least 2 GB free beyond the weights for the KV cache and buffers, more if you work with long documents. The table suggests a starting point by use case, based on the weights in the QuelLLM catalog.
| Usage | Starting model | Q4 weight | Headroom on 16 GB |
|---|---|---|---|
| General chat and writing | Gemma 4 12B | 7 GB | 9 GB, long context possible |
| Code, development agent | Devstral Small 2 24B | 14 GB | 2 GB, short context |
| Fast reasoning (MoE) | gpt-oss 20B | 13 GB | 3 GB |
| RAG with long prompts | Qwen 3.5 9B | 6 GB | 10 GB for context |
| Lightweight tasks running continuously | Granite 4.2 8B | 4.6 GB | 11 GB |
For RAG, the 9070 XT’s matrix computation is a real advantage: a prompt containing several thousand tokens is processed quickly, and a 9B model in Q4 leaves 10 GB for the KV cache, enabling a long context. For coding, a 14 GB model fits but with a short context; you then need to quantize the KV cache or reduce the context window.
#Context and KV cache: the two settings that matter
The Ollama documentation indicates a default context window of 4,096 tokens, which the OLLAMA_CONTEXT_LENGTH variable changes. The KV cache is quantized with OLLAMA_KV_CACHE_TYPE, f16 by default, when Flash Attention is active; Ollama uses it automatically when the backend and card support it. Increase the context gradually while monitoring ollama ps: as soon as part of the model spills onto the processor, step back.
#FP8 on RX 9070 XT: what it really changes
AMD lists 389 TFLOPs of matrix FP8 compute for the 9070 XT, which is not listed on the RX 7900 XTX specifications. In practice, with Ollama and Q4 or Q5 GGUF models, you do not use FP8: the weights are quantized integers. FP8 is mainly used by engines that load FP8 weights, such as vLLM, which lists the RX 9000. Treat it as headroom for later, not an immediate gain, and do not buy the card for it: the 16 GB capacity and bandwidth remain the two criteria that determine what you can run today.
#RX 9070 XT or RTX 5070 Ti
NVIDIA describes the RTX 5070 Ti with 16 GB of GDDR7 on a 256-bit bus: the same capacity as the AMD card, with memory from a different generation. At equal capacity, the models that fit are the same; the differences are generation speed, the CUDA ecosystem, and price, which we do not lock in here. Check the price tracker for local AI cards: it is usually the deciding factor.
| Your situation | Preferred card | Reason |
|---|---|---|
| Linux, Ollama, or LM Studio, measured budget | RX 9070 XT | ROCm 7 supported, 16 GB, FP8 compute listed |
| Windows and no desire to debug | RTX 5070 Ti | CUDA ecosystem, more universal drivers |
| Projects requiring CUDA (some fine-tuning pipelines) | RTX 5070 Ti | CUDA remains the default target for many projects |
| RAG with long prompts | RX 9070 XT or RTX 5070 Ti | Compute matters as much as bandwidth |
| 27-billion-parameter models and larger | 24 GB card | No 16 GB option is enough |
#Power supply and case
AMD recommends a power supply of at least 750 W and lists 304 W of typical board power, with two 8-pin connectors. Add the power consumption of the processor, drives, and fans to size the system, and check your manufacturer's model length and thickness: these cards are often large. A well-ventilated case limits noise and clock-speed drops under sustained load, which matters when the card generates text for hours. For an AI server that stays powered on, idle and load power consumption matter just as much as the purchase price over time.
- Which LLM for 16 GB of VRAM: the complete list
- Ollama with AMD GPU (ROCm): installation
- LLM on RX 7900 XTX (24 GB)
- LLM on RX 9060 XT (8 / 16 GB)
- Hardware profile RX 9070 XT
- Local AI graphics card prices
- Source: AMD's Radeon RX 9070 XT specification sheet
- Source: Ollama GPU documentation
- Source: LM Studio 0.3.19, support for AMD 9000 series
- Source: llama.cpp performance table on ROCm
#Frequently asked questions
Is the RX 9070 XT good for local LLMs?+
How do you use LM Studio with a RX 9070 XT?+
Do you need ROCm or Vulkan on RX 9070 XT?+
How many tokens per second does a RX 9070 XT produce?+
Does RX 9070 XT support FP8?+
What power supply should you choose for a RX 9070 XT?+
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.