Beginner 11 minRadeon RX 9000

Which LLM on Radeon RX 9060 XT (8 / 16 GB) ?

Direct response

The RX 9060 XT can be used for local LLMs, provided you choose the 16 GB version: the same 320 GB/s as the 8 GB version according to AMD, but enough memory for 12- to 24-billion-parameter models in Q4. The 8 GB version is limited to 9-billion-parameter models. It is supported by ROCm 7 on Linux.

The RX 9060 XT is the entry-level Radeon RDNA 4 for local AI. This guide compares its two memory versions, explains what fits in each, gives the speed ceiling allowed by its 320 GB/s, and details its Linux and Windows compatibility.

Choosing a machine? Our picks by budget →

By Mohamed Meguedmi·Update 2026-09-29·Tested on Windows, macOS, and Linux
Recommended hardware

Buying alternative for this guide: Radeon RX 9070 XT 16GB (ASUS Prime OC).

Why this choice? Our complete guide on Radeon RX 9070 XT 16GB (ASUS Prime OC) →

Compare all options by budget, from €800 to €3,500 →

On the go: which laptop for local AI →

Affiliate links — commission possible at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.

#RX 9060 XT: the 16 GB version is the only serious option for LLMs

To run LLMs locally, choose the Radeon RX 9060 XT in the 16 GB version. Both versions share the same GPU and the same 320 GB/s bandwidth; only capacity changes, and that is what determines what you can load. With 8 GB, you are limited to 7- to 9-billion-parameter models in Q4, with a short context. With 16 GB, you can move to 12- to 24-billion-parameter models, a longer context, and more faithful quantization. At launch, AMD listed $299 for the 8 GB version and $349 for the 16 GB version: a $50 difference to double the memory.

RX 9060 XT 8 GB and 16 GB according to AMD
Feature8 GB16 GB
Memory8 GB GDDR616 GB GDDR6
Memory bus and bandwidth128-bit, up to 320 GB/s128-bit, up to 320 GB/s
Typical card power150 W160 W
Recommended minimum power supply450 W450 W
U.S. list price at launch (May 2025)299 $349 $

Prices change; to find today's price and the price per gigabyte of VRAM, check the local AI graphics card price tracker. The key principle: on an inexpensive card, the premium for the 16 GB version is small compared with what it unlocks.

#ROCm, Vulkan, and Windows compatibility

The RX 9060 XT appears in the ROCm 7.2.1 compatibility matrix for Radeon cards, for Ubuntu 22.04, 24.04, and 25.10. Ollama also lists it among the cards supported by ROCm on Linux, with the ROCm v7 driver. On Linux, installing Ollama with ROCm therefore follows the same path as with other Radeon cards: the Ollama guide with AMD GPUs details the process.

Be cautious on Windows. The list of cards that the Ollama documentation gives for ROCm on Windows stops at RX 7000 and does not mention RX 9000. Ollama also states that Vulkan provides additional support on Windows and Linux, enabled by default. A RX 9060 XT may therefore work on Windows through Vulkan; check with ollama ps that the model is actually on the GPU and not the CPU. LM Studio and llama.cpp also offer a Vulkan backend.

i
One point to verify yourself
The site does not test cards: the compatibility listed above is what AMD and Ollama advertise. After installation, two commands are enough to confirm it: rocm-smi on Linux to view the card, and ollama ps after a response to read the percentage running on the GPU.

#What fits in 8 GB and 16 GB

The weights below come from the QuelLLM catalog, in Q4 unless otherwise noted. In addition to these weights, account for the KV cache, which grows with the context, and roughly one gigabyte of headroom for the system and buffers. A model that fills the card to 100% with weights is therefore too large.

Model weights (QuelLLM catalog) and verdict by version
ModelWeights8 GB16 GB
Qwen 3.5 9B Q46 GBFair, short contextComfortable
Qwen 3.5 9B Q810 GBNoComfortable
Gemma 4 12B Q47 GBToo tightComfortable
Gemma 4 12B Q813 GBNoTight
gpt-oss 20B Q413 GBNoYes, moderate context
Mistral Small 24B Q414 GBNoTight, 2 GB for context
Qwen 3.8 27B Q416 GBNoNo, it overflows

The practical conclusion fits in one sentence: 16 GB opens up the 12- to 24-billion-parameter category, where you’ll find significantly more capable models for reasoning, code, and French. The guide to 16 GB of VRAM provides the complete list; the 8 GB guide covers tricks for stretching memory.

#Expected speed: bandwidth sets the ceiling

Text generation reads most of the model's weights at each token, so maximum speed is bandwidth divided by the size of the weights read. At 320 GB/s, that's a ceiling of about 53 tokens per second for a Qwen 3.5 9B in Q4 (6 GB), 46 for a Gemma 4 12B (7 GB), 23 for a Mistral Small 24B (14 GB), and about 32 for the same Qwen 3.5 9B in Q8 (10 GB). These are theoretical ceilings, never reached exactly; MoE models, which read only their active weights, partly avoid this limitation.

The only comparable public benchmark is the llama.cpp community’s ROCm table: for the RX 9060 XT, one participant reports 1 420 tokens per second for prompt processing and 68 for generation on a Llama 2 7B in Q4_0, without Flash Attention. This small model is much lighter than those in the previous table. It provides an order of magnitude, not the speed of a 12B or 24B on your setup, which will vary with the driver, operating system, and manufacturer’s card.

Theoretical ceiling at 320 GB/s (bandwidth ÷ catalog Q4 weights)
ModelWeights read per tokenTheoretical ceiling
Qwen 3.5 9B Q46 GB≈ 53 tokens/s
Qwen 3.5 9B Q810 GB≈ 32 tokens/s
Gemma 4 12B Q47 GB≈ 46 tokens/s
Mistral Small 24B Q414 GB≈ 23 tokens/s
→
Q8 or Q4: the real cost
Moving from Q4 to Q8 nearly doubles the weights read per token, so it roughly halves the speed ceiling. On a 320 GB/s card, Q4_K_M remains the best compromise for chat; reserve Q8 for tasks where accuracy matters more than responsiveness.

#Set Ollama so you don't exceed 16 GB

On 16 GB, every gigabyte matters. Three settings prevent silent overflow onto the CPU, which causes a speed drop without an error message. The first is context size: Ollama’s documentation specifies a default window of 4,096 tokens, adjustable with OLLAMA_CONTEXT_LENGTH; each doubling of the context increases the KV cache. The second is the KV cache itself, quantifiable with OLLAMA_KV_CACHE_TYPE (f16 by default, q8_0 to reduce it) when Flash Attention is active. The third is monitoring: after a response, ollama ps displays the split between GPU and CPU.

Server with 16k context and q8_0 KV cache
OLLAMA_CONTEXT_LENGTH=16384 OLLAMA_KV_CACHE_TYPE=q8_0 ollama serve
# dans un autre terminal
ollama run <nom-du-modèle>
ollama ps

One example of the reasoning: a 14 GB model such as Mistral Small 24B in Q4 on a 16 GB card leaves about 2 GB for context, cache, and buffers. If ollama ps shows a GPU/CPU split, reduce the context first, then quantize the KV cache, then drop down one quantization level, rather than letting the model spill over. The guide to KV cache quantization explains the process in detail.

#Buy, wait, or target 24 GB

Three situations arise. If you already have a recent desktop PC and a 450 W or higher power supply, the 9060 XT 16 GB is a straightforward way to move from limited use with small models to serious use for chat, translation, and light coding. If you are targeting models with 27 to 35 billion parameters, neither 16 GB nor its 320 GB/s is sufficient: look at 24 GB cards, covered in the RX 7900 XTX guide. If you cannot decide based on price, compare the cost per gigabyte of VRAM in the site’s tracker, which is the right criterion for local AI.

#Two RX 9060 XT: 32 GB for less?

The question keeps coming up in searches. In llama.cpp, the default distribution mode splits the model by layers across the cards in a pipeline, and an option sets the proportion assigned to each card. Two 9060 XT 16 GB cards provide 32 GB of total VRAM at about 320 W of typical power according to AMD, which is enough to load a Qwen 3.8 27B in Q4 (16 GB of weights) with plenty of context, or a 30- to 35-billion-parameter MoE (19 to 21 GB).

Two limitations to know. Per-token speed doesn't double: the cards work one after another on each token, and each retains its 320 GB/s of bandwidth. The motherboard must also provide two usable PCIe slots and a well-ventilated case. A single 24 GB card, such as the 7900 XTX, offers 960 GB/s over a single bus; for a model that fits on it, it generates noticeably faster.

#FP8: an advantage on paper, few consequences with Ollama

The Radeon RX 9000 cards use the RDNA 4 generation, and AMD lists 205 TFLOPs of FP8 compute for the 9060 XT, which the RX 7900 XTX specifications do not mention. For typical use with Ollama or LM Studio, this matters little: GGUF models are quantized as integers (Q4, Q5, Q8), not FP8. It becomes useful with engines that use FP8 weights, such as vLLM, whose documentation lists the RX 9000. Do not use it as a purchasing argument if you are sticking with Ollama.

#RX 9060 XT or RTX 5060 Ti: the 16 GB showdown

NVIDIA offers the RTX 5060 Ti with 16 GB and 8 GB, using GDDR7 on a 128-bit bus, versus GDDR6 on 128 bits for the AMD card. The faster memory gives NVIDIA the bandwidth advantage, and therefore higher generation speed; see the NVIDIA spec sheet for the exact throughput. The other advantage is software: CUDA is the default target for most AI projects.

Decision criteria at 16 GB
Your priorityChoiceReason
Maximum generation speed, CUDA requiredRTX 5060 Ti 16 GBGDDR7 memory, CUDA ecosystem
Linux, Ollama or llama.cpp, on a measured budgetRX 9060 XT 16 GBROCm 7 and Vulkan cover common workloads
Windows, no desire to debugRTX 5060 Ti 16 GBMost universal drivers and tools
Models larger than 16 GBNeitherAim for 24 GB: RX 7900 XTX or RTX 3090
Very tight budgetRX 9060 XT 8 GBReserved for models with 9 billion parameters
!
No purchasing advice without today's price
Price differences between these cards change every week and can reverse the decision. Compare them on the site's price tracker before deciding.

#Frequently asked questions

Frequently asked questions
RX 9060 XT 8 GB or 16 GB for LLMs?+
Choose the 16 GB model. Both have 320 GB/s of bandwidth according to AMD, but the 8 GB version is limited to 9-billion-parameter models in Q4 with a short context. The 16 GB version loads Gemma 4 12B, gpt-oss 20B, or Mistral Small 24B in Q4. At launch, the suggested price difference was 50 dollars.
How many tokens per second does a RX 9060 XT produce?+
It depends on the model. The theoretical ceiling is bandwidth (320 GB/s) divided by the weights: about 53 tokens/s for a 9B in Q4 and 23 for a 24B. On a Llama 2 7B in Q4_0, a participant in the llama.cpp leaderboard records 68 tokens/s during generation on ROCm.
Does RX 9060 XT work with Ollama?+
Yes on Linux: Ollama lists the cards supported by ROCm v7. On Windows, its ROCm list stops at RX 7000, but Vulkan, enabled by default, provides additional support. Check with ollama ps that the model is placed 100% on the GPU.
Does RX 9060 XT support FP8?+
The AMD spec sheet lists 205 TFLOPs of FP8 matrix compute, which does not appear on the RX 7900 XTX spec sheet. Ollama and Q4 GGUF models do not use it directly. FP8 is mainly relevant to vLLM, which lists Radeon RX 9000 among its supported hardware.
Can you connect two RX 9060 XT to get 32 GB?+
Yes, with llama.cpp, which distributes the layers across the GPUs by default. You get 32 GB of capacity, but not twice the speed: each card retains its 320 GB/s and they work in turn. According to AMD, plan on about 320 W for both cards, plus the rest of the PC.
What power supply do you need for a RX 9060 XT?+
AMD recommends at least 450 W for both versions, with typical power of 150 W for the 8 GB model and 160 W for the 16 GB model. Add the power draw of the processor and drives when sizing the power supply, then check the specifications for your exact model, since some card manufacturers recommend more than AMD's minimum.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.

Prices in euros (€) are French market prices including VAT, as checked by BestLLMfor. US prices differ: the Amazon buttons show the current US price.