Intermediate 11 minRadeon RX 9000

Which LLM on Radeon RX 9070 XT (16 GB) ?

Direct response

The RX 9070 XT (16 GB GDDR6, up to 640 GB/s according to AMD) is a good card for local LLMs on Linux with ROCm 7, or through Vulkan on Windows. It loads models with up to 13–14 GB of Q4 weights without offloading, such as gpt-oss 20B or Mistral Small 24B. Dense models with 27 billion parameters are out of reach.

The RX 9070 XT is AMD's high-end RDNA 4 Radeon. This guide explains what its 16 GB enables, how to install it on Linux or Windows, what public benchmarks show, and when a RTX 5070 Ti or a 24 GB card is preferable.

Choosing a machine? Our picks by budget →

By Mohamed Meguedmi·Update 2026-09-29·Tested on Windows, macOS, and Linux
Recommended hardware

For this setup: Radeon RX 9070 XT 16GB (ASUS Prime OC).

Why this choice? Our complete guide on Radeon RX 9070 XT 16GB (ASUS Prime OC) →

Compare all options by budget, from €800 to €3,500 →

On the go: which laptop for local AI →

Affiliate links — possible commission at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.

#RX 9070 XT for LLMs: 16 GB, 640 GB/s, and strong compute

The Radeon RX 9070 XT is the most capable consumer Radeon for LLMs among the RX 9000: 16 GB of GDDR6 on a 256-bit bus, bandwidth of up to 640 GB/s according to AMD, and substantially greater matrix compute than the RX 7900 XTX in FP16 and FP8. It runs models whose weights remain below about 13 to 14 GB in Q4 without overflow: a 9B or 12B with a long context, gpt-oss 20B, Mistral Small 24B with a short context. Its limitation is capacity: dense models with 27 billion parameters in Q4 (about 16 GB) do not fit.

RX 9070 XT technical specifications (source: AMD)
FeatureValue
Memory16 GB GDDR6, 256-bit bus
Bandwidthup to 640 GB/s
Stream processors / compute units4 096 / 64
AI accelerators128
FP16 / FP8 matrix multiplication195 / 389 TFLOPs
Typical card power304 W
Recommended minimum power supply750 W
Power connectors2 x 8-pin
i
Bandwidth: 640 GB/s, not 644
Some online specifications, including an earlier version of this guide, list 644 GB/s. AMD's official specification says “up to 640 GB/s”; that is the figure to use in your calculations.

#Software: ROCm, Vulkan, LM Studio, vLLM

The RX 9070 XT appears in AMD's Radeon documentation matrix for ROCm 7.2.1 (Ubuntu 22.04, 24.04, and 25.10). Ollama lists it among the cards supported on Linux with the ROCm v7 driver and maps the gfx1201 target to the 9070 XT in its LLVM targets table. LM Studio announced support for AMD 9000-series GPUs on Linux with ROCm in version 0.3.19, released in July 2025. vLLM lists the Radeon RX 9000 (gfx1200 and gfx1201) among its AMD hardware.

Under Windows, the ROCm list from Ollama stops at the RX 7000: the RX 9000 are not included. Ollama nevertheless specifies that Vulkan provides additional support under Windows and Linux, enabled by default. A 9070 XT under Windows therefore effectively uses Vulkan; check with ollama ps that the model is actually on the GPU, and consult the llama.cpp Vulkan guide if you compile it yourself.

#Getting started on Linux with Ollama

  1. 01
    Install the ROCm driver
    Follow AMD’s ROCm documentation with the amdgpu-install utility, as Ollama requires, add your user to the render and video groups, then reboot.
  2. 02
    Confirm detection
    rocminfo must list a GPU agent and rocm-smi must display 16 GB. Without that, Ollama will use the processor.
  3. 03
    Run your first model
    Install Ollama with its official script and start a model that fits in 16 GB, such as a 9B in Q4.
  4. 04
    Check placement
    After the first response, ollama ps should show 100% GPU. Shared GPU and processor usage indicates that the model or context is too large.
Checks
rocminfo | grep -i gfx
rocm-smi --showmeminfo vram
ollama ps

#Which models fit in 16 GB

The weights come from the QuelLLM catalog, in Q4 unless noted otherwise. Add the KV cache and about one gigabyte of headroom. The full list by tier is in the 16 GB VRAM guide.

Weights (QuelLLM catalog) and theoretical ceiling at 640 GB/s
ModelWeightsDoes it fit in 16 GB?Theoretical ceiling
Qwen 3.5 9B Q46 GBYes, very broad≈ 107 tokens/s
Qwen 3.5 9B Q810 GBYes64 tokens/s
Gemma 4 12B Q47 GBYes, long context≈ 91 tokens/s
Gemma 4 12B Q813 GBYes, short context≈ 49 tokens/s
gpt-oss 20B (MoE)13 GBYes, moderate contextdepends on the active weights
Mistral Small 24B Q414 GBTight≈ 46 tokens/s
Qwen 3.8 27B Q416 GBNo-

The ceiling is bandwidth divided by the weights read at each token; no card reaches it exactly. MoE models such as gpt-oss 20B read fewer weights per token than their total size and therefore run faster than a dense model of the same size, provided they fit entirely in VRAM.

!
What “16 GB” doesn't allow
In Q4, models with 27 to 35 billion parameters weigh 16 to 21 GB according to the catalog: they spill into RAM, and speed collapses. For this category, see the 24 GB guide or the RX 7900 XTX.

#What the public measurements show

The llama.cpp community table on ROCm measures the same model, Llama 2 7B in Q4_0, on several cards. For the RX 9070 XT, one participant reports 5 055 tokens per second when reading the prompt (pp512) and 101 during generation (tg128), without Flash Attention. On the 7900 XTX, the table gives 3 552 and 167; on the 9060 XT, 1 420 and 68.

Llama 2 7B Q4_0, llama.cpp ROCm (community table, without Flash Attention)
CardAMD bandwidthPrompt processingGeneration
RX 7900 XTX960 GB/s3 552 t/s167 t/s
RX 9070 XT640 GB/s5 055 t/s101 t/s
RX 9060 XT320 GB/s1 420 t/s68 t/s

Two takeaways. First, the 9070 XT reads a long prompt about 40% faster than the 7900 XTX, despite lower bandwidth: that's the advantage of more powerful matrix computation, which is useful for RAG, long documents, and coding agents that return a lot of context. Second, during generation, the 7900 XTX remains ahead, as its 50% higher bandwidth would suggest.

i
A community table is not a benchmark
In this same table, the RX 9070 (non-XT) shows 114 tokens/s during generation, higher than the 9070 XT at 101, even though it is supposed to be slower. The rows come from different versions of llama.cpp and different drivers. Treat these figures as orders of magnitude, never as a precise ranking.

#Which model to choose for each use case

With 16 GB, the right choice depends less on the maximum size than on the headroom left for context. A simple rule: keep at least 2 GB free beyond the weights for the KV cache and buffers, more if you work with long documents. The table suggests a starting point by use case, based on the weights in the QuelLLM catalog.

Usage starting point on 16 GB
UsageStarting modelQ4 weightHeadroom on 16 GB
General chat and writingGemma 4 12B7 GB9 GB, long context possible
Code, development agentDevstral Small 2 24B14 GB2 GB, short context
Fast reasoning (MoE)gpt-oss 20B13 GB3 GB
RAG with long promptsQwen 3.5 9B6 GB10 GB for context
Lightweight tasks running continuouslyGranite 4.2 8B4.6 GB11 GB

For RAG, the 9070 XT’s matrix computation is a real advantage: a prompt containing several thousand tokens is processed quickly, and a 9B model in Q4 leaves 10 GB for the KV cache, enabling a long context. For coding, a 14 GB model fits but with a short context; you then need to quantize the KV cache or reduce the context window.

#Context and KV cache: the two settings that matter

The Ollama documentation indicates a default context window of 4,096 tokens, which the OLLAMA_CONTEXT_LENGTH variable changes. The KV cache is quantized with OLLAMA_KV_CACHE_TYPE, f16 by default, when Flash Attention is active; Ollama uses it automatically when the backend and card support it. Increase the context gradually while monitoring ollama ps: as soon as part of the model spills onto the processor, step back.

32k context with q8_0 KV cache
OLLAMA_CONTEXT_LENGTH=32768 OLLAMA_KV_CACHE_TYPE=q8_0 ollama serve

#FP8 on RX 9070 XT: what it really changes

AMD lists 389 TFLOPs of matrix FP8 compute for the 9070 XT, which is not listed on the RX 7900 XTX specifications. In practice, with Ollama and Q4 or Q5 GGUF models, you do not use FP8: the weights are quantized integers. FP8 is mainly used by engines that load FP8 weights, such as vLLM, which lists the RX 9000. Treat it as headroom for later, not an immediate gain, and do not buy the card for it: the 16 GB capacity and bandwidth remain the two criteria that determine what you can run today.

#RX 9070 XT or RTX 5070 Ti

NVIDIA describes the RTX 5070 Ti with 16 GB of GDDR7 on a 256-bit bus: the same capacity as the AMD card, with memory from a different generation. At equal capacity, the models that fit are the same; the differences are generation speed, the CUDA ecosystem, and price, which we do not lock in here. Check the price tracker for local AI cards: it is usually the deciding factor.

How to decide at 16 GB
Your situationPreferred cardReason
Linux, Ollama, or LM Studio, measured budgetRX 9070 XTROCm 7 supported, 16 GB, FP8 compute listed
Windows and no desire to debugRTX 5070 TiCUDA ecosystem, more universal drivers
Projects requiring CUDA (some fine-tuning pipelines)RTX 5070 TiCUDA remains the default target for many projects
RAG with long promptsRX 9070 XT or RTX 5070 TiCompute matters as much as bandwidth
27-billion-parameter models and larger24 GB cardNo 16 GB option is enough

#Power supply and case

AMD recommends a power supply of at least 750 W and lists 304 W of typical board power, with two 8-pin connectors. Add the power consumption of the processor, drives, and fans to size the system, and check your manufacturer's model length and thickness: these cards are often large. A well-ventilated case limits noise and clock-speed drops under sustained load, which matters when the card generates text for hours. For an AI server that stays powered on, idle and load power consumption matter just as much as the purchase price over time.

#Frequently asked questions

Frequently asked questions
Is the RX 9070 XT good for local LLMs?+
Yes, for models that fit in its 16 GB: a 9B or 12B with a long context, gpt-oss 20B, Mistral Small 24B with a short context. Its 640 GB/s bandwidth and matrix compute make it fast, particularly for prompt processing. It is not suitable for dense models with 27 billion parameters.
How do you use LM Studio with a RX 9070 XT?+
LM Studio has announced support, since version 0.3.19, for AMD 9000-series GPUs on Linux with ROCm. On Windows, the Vulkan engine in llama.cpp is the fallback option. In both cases, load a model under 14 GB and check in the interface that all layers are on the GPU.
Do you need ROCm or Vulkan on RX 9070 XT?+
On Linux, ROCm 7 is the official path for Ollama, with the gfx1201 target for this card. On Windows, the ROCm list for Ollama does not mention RX 9000, so Vulkan, which is enabled by default, is the practical option. Compare both with llama-bench on your model before deciding.
How many tokens per second does a RX 9070 XT produce?+
It depends on the model. On a Llama 2 7B in Q4_0, llama.cpp’s community table reports 101 tokens/s for generation on ROCm. For another model, the theoretical ceiling is 640 GB/s divided by the weights: about 46 tokens/s for a 24B model weighing 14 GB, before any losses.
Does RX 9070 XT support FP8?+
On paper, yes: AMD lists 389 TFLOPs of FP8 matrix compute. With Ollama and Q4 GGUF models, you don’t benefit from it directly because the weights are quantized integers. FP8 is used by engines that load FP8 weights, such as vLLM, which lists the RX 9000.
What power supply should you choose for a RX 9070 XT?+
AMD specifies a minimum of 750 W and typical power consumption of 304 W, with two 8-pin power connectors. Add the power consumption of the rest of the PC and leave some headroom, especially if the processor is powerful. Also check your manufacturer's model specifications, which may recommend more.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.