Intermediate 11 minRadeon RX 7000

Which LLM on Radeon RX 7900 XT (20 GB) ?

Direct response

The Radeon RX 7900 XT (20 GB GDDR6, 800 GB/s, 315 W) loads 24B to 26B models in Q4 without spilling, with a comfortable context, and can still handle a 30B model at the cost of a short context. It's memory, not speed, that sets it apart: on Llama 2 7B, public benchmarks report 123 tokens per second under Vulkan, barely faster than a RX 7800 XT with 16 GB, and far behind the RX 7900 XTX.

Twenty gigabytes of VRAM is rare capacity on a consumer card, and it changes what you can run: it’s the difference between a cramped 24B and a 24B with context. This guide separates what those 20 GB actually provide from what 800 GB/s of bandwidth might suggest, using cited public measurements and correcting misconceptions about this card.

Choosing a machine? Our picks by budget →

By Mohamed Meguedmi·Update 2026-09-30·Tested on Windows, macOS, and Linux
Recommended hardware

Buying alternative for this guide: Radeon RX 9070 XT 16GB (ASUS Prime OC).

Why this choice? Our complete guide on Radeon RX 9070 XT 16GB (ASUS Prime OC) →

Compare all options by budget, from €800 to €3,500 →

On the go: which laptop for local AI →

Affiliate links — possible commission at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.

#What a RX 7900 XT does in 2026

The RX 7900 XT is for anyone who wants 24- to 26-billion-parameter models without constraining the context. Its 20 GB accommodates Devstral Small 2 24B (about 14 GB in Q4) with 6 GB to spare, versus 1 GB on a 16 GB card, and the MoE Gemma 4 26B-A4B (about 16 GB) with 4 GB to spare. In speed, however, it disappoints anyone counting on its 800 GB/s: 123.18 tokens per second on the reference 7B under Vulkan, versus 118.27 for a RX 7800 XT whose bandwidth is 22% lower. AMD officially lists it in ROCm, and Ollama supports it on Linux and Windows.

Architecture
RDNA 3, Navi 31 with one GCD and five MCDs, launched on December 13, 2022.
Compute
5,376 stream processors, 84 compute units.
Memory
20 GB GDDR6, 320-bit bus, 800 GB/s.
Power consumption
315 W of board power: check the power supply recommended by your model’s manufacturer.
Pricing
Not specified here: see /prix-gpu-ia.

#The 20 GB memory budget

Allow about 19 GB for the model and its KV cache when the card is not driving the display. The table subtracts the Q4 size from the QuelLLM catalog from this budget to show the headroom available for context. This headroom is the real criterion: a model that “fits” with 1 GB left has room for only a few thousand tokens of context.

What a 7900 XT supports in Q4 (weights only, 19 GB budget)
ModelWeightsHeadroom for contextVerdict
gpt-oss 20B≈ 13 GB≈ 6 GBVery comfortable, long context
Devstral Small 2 24B≈ 14 GB≈ 5 GBComfortable, long context possible
Gemma 4 26B-A4B (MoE)≈ 16 GB≈ 3 GBFits, medium context
Granite 4.1 30B≈ 17 GB≈ 2 GBFits, short context
Gemma 4 31B≈ 18 GB≈ 1 GBCramped, very short context

Mixture-of-experts models deserve a note. Only a fraction of an MoE's parameters work on each token, making generation faster than with a dense model of the same size, but all the weights still have to fit in memory. The 26B-A4B is the kind of model where 20 GB makes the difference: too large for 16 GB with context, comfortable in 20. For even larger models, such as the Qwen3.6 35B-A3B family, consult the dedicated guide rather than guessing: the exact size depends on the quantization used.

To convert the margin into a number of context tokens, you need the KV-cache cost per token, which depends on the model architecture: number of layers, attention heads, and cache precision. So there is no universal tokens-per-gigabyte rule. The reliable method is measurement: load the model with a modest context, note the VRAM usage with rocm-smi, increase the window, and repeat. The site's VRAM calculator automates this estimate for models in the catalog, and quantizing the KV cache to 8-bit roughly halves the context's share.

#Measured throughput: bandwidth is not everything

The figures come from public discussions in the llama.cpp repository, which use Llama 2 7B in Q4_0 and llama-bench; they are not measurements from the site. Under Vulkan, the RX 7900 XT reads a 512-token prompt at 2,941.58 tokens per second and generates at 123.18. With Flash Attention, the figures are 2,701.13 and 120.62. Under ROCm, the discussion reports 3,098.38 for reading and 116.15 for generation. The two sets converge: generation tops out at around 120 tokens per second.

The theoretical bandwidth ceiling, 800 GB/s divided by 3.82 GB of weights, is about 209 tokens per second; the card reaches only 59% under Vulkan. This is the lowest efficiency among the AMD cards in this series of measurements (83% for a RX 6700 XT, 77% for a RX 7900 GRE, 72% for a RX 7800 XT). The likely explanation is that a 7-billion-parameter model is too small to saturate the memory of a large chip; better efficiency can be expected on larger models, but none of the measurements in these series proves it, and it would be imprudent to announce throughput figures for a 24B.

RX 7900 XT on Llama 2 7B Q4_0 (public llama.cpp measurements)
BackendFlash Attentionpp512 processingGeneration tg128
Vulkanno2 941,58 t/s123,18 t/s
Vulkanyes2 701,13 t/s120,62 t/s
ROCmno3 098,38 t/s116,15 t/s

For a ballpark estimate on large models, a cautious bound is to apply the measured 59% efficiency to the theoretical ceiling: a 24B model in Q4 at 14 GB gives 800 divided by 14, or about 57 tokens per second at the ceiling, and therefore roughly 34 tokens per second. This is a calculation, not a measurement; if efficiency is better on a large model, the actual throughput will be higher.

#Facing RX 7900 XTX: the real gap is much larger than people say

You often read that the 7900 XT is “10 to 12% slower” than the 7900 XTX. Public benchmarks tell a different story. On the reference 7B model under Vulkan, the XTX generates 182.63 tokens per second, versus 123.18 for the XT: 48% more. The stream processor ratio (6,144 versus 5,376) isn’t enough to explain it; memory bandwidth and the chip itself create the gap. What the XT preserves is the entry price: with 20 GB, it covers most of what 24B to 26B models require.

RX 7900 XT and RX 7900 XTX (Llama 2 7B Q4_0, public measurements)
MetricRX 7900 XTRX 7900 XTX
VRAM20 GB24 GB
Vulkan, tg128 generation123,18 t/s182,63 t/s
Vulkan, pp512 prompt processing2 941,58 t/s3 726,99 t/s

This correction changes the trade-off: if generation speed matters to you, the 48% gap outweighs the additional 4 GB of VRAM, whereas if you mainly want capacity for context, the XT delivers most of what the XTX offers for a 24B model. On a 30B model or larger, however, the XTX's 24 GB eliminates context compromises, and that's where its price premium is justified.

#Versus the RTX 4070 Ti Super: more memory, less speed

The RTX 4070 Ti Super offers 16 GB, which is 4 less. On the reference 7B model under Vulkan, it generates at 129.45 tokens per second, 5% better than the RX 7900 XT, and processes the prompt at 6,099 tokens per second, twice as fast. So the XT is not “faster,” contrary to what you sometimes read: it is larger. Its advantage is capacity, which enables 24B to 26B models with context.

RX 7900 XT and RTX 4070 Ti Great under Vulkan (Llama 2 7B Q4_0, without FA)
CriterionRX 7900 XTRTX 4070 Ti Super
VRAM20 GB16 GB
pp512 processing2 941,58 t/s6 099,18 t/s
Generation tg128123,18 t/s129,45 t/s

#When a model exceeds 20 GB

A model that’s too large doesn’t cause an error: Ollama and llama.cpp place as many layers as possible on the card and leave the rest to the CPU. In llama.cpp, the -ngl option sets the number of layers sent to the GPU, and the value 100 used in the measurements above means “all.” Reducing this value frees up VRAM, at the cost of slower generation, since each token must then pass through system memory, which is much slower than GDDR6. The rule of thumb: as long as ollama ps shows 100% GPU, you’re in the fast zone; as soon as a proportion appears on the CPU side, speed drops in steps.

#20 GB, 16 GB, or 24 GB: how to choose

What capacity for which use
Intended use16 GB20 GB (RX 7900 XT)24 GB
Chat with an 8B to 14B modelSufficientUnnecessaryUnnecessary
24B coding assistant, long contextCrampedComfortableComfortable
26B MoE with medium contextToo close to callSufficientSufficient
30B to 35B model with long contextImpossibleJustRequired

The principle: pay for the capacity your largest model requires, not for what you might use someday. If your target is a 24B model with context, 20 GB is the first tier where it becomes comfortable; beyond that, each additional gigabyte must be justified by a specific model.

#Getting started

  1. 01
    Choose the system
    AMD supports Radeon RX 7000 only on Ubuntu 24.04.4, Ubuntu 22.04.5, RHEL 10.1, and RHEL 9.7. On Windows, Ollama uses a ROCm v7 or HIP7 stack; Vulkan serves as a fallback.
  2. 02
    Install the driver and Ollama
    On Linux, Ollama requires the ROCm v7 driver. Install it with the amdgpu-install utility, then Ollama.
  3. 03
    Load a 24B model
    Choose a Devstral Small 2 24B in Q4_K_M, run it, and check with ollama ps that it is using the GPU at 100%.
  4. 04
    Adjust the context
    Increase the context window in stages while monitoring VRAM: the goal is to stay below 19 GB total.
  5. 05
    Compare Vulkan and ROCm
    Measure with llama-bench on your model. The two backends are close on this card in publicly available benchmarks.
Terminal
ollama ps
rocm-smi --showmeminfo vram
./llama-bench -m llama-2-7b.Q4_0.gguf -ngl 100 -fa 0,1

#2026 verdict

A good choice
If you want to run 24B to 26B models with context and the cost per GB of VRAM on /prix-gpu-ia is favorable.
A middle-ground choice
If speed comes first: generation tops out near the level of a 16 GB card, and the RTX 4070 Ti Super reads prompts twice as fast.
An insufficient choice
For 30B to 35B models with a long context, where 24 GB becomes necessary.

#Frequently asked questions

FAQ
Is the RX 7900 XT good for local LLMs?+
Yes, for its memory: 20 GB lets you load a 24B model in Q4 with a comfortable context, which 16 GB can barely handle. Its generation speed, about 123 tokens per second on a 7B in Q4_0 under Vulkan, remains close to that of a RX 7800 XT.
RX 7900 XT or RX 7900 XTX for local AI?+
The XTX provides 24 GB and 48% faster generation on the reference 7B under Vulkan (182.63 versus 123.18 tokens per second). The XT is sufficient for 24B to 26B models and costs less. Compare current prices on the site's dedicated page.
Is 20 GB of VRAM enough for a 30B model?+
A 30B model such as Granite 4.1 30B weighs about 17 GB in Q4: it fits with a short context, leaving about 2 GB of headroom. For a long context or 35B models such as Qwen3.6 35B-A3B, 24 GB becomes necessary. See the dedicated guide for this model.
How many tokens per second does a RX 7900 XT produce?+
On Llama 2 7B in Q4_0, public discussions in the llama.cpp repository report 123.18 tokens per second under Vulkan and 116.15 under ROCm during generation. For a 24B in Q4, the calculation gives approximately 34 tokens per second: an estimate to measure on your system.
RX 7900 XT or RTX 4070 Ti Great for an LLM?+
The XT for capacity (20 GB vs. 16): it loads 24B to 26B models with context. The RTX 4070 Ti Super generates 5% faster on the reference 7B and reads prompts twice as fast, with the CUDA ecosystem. The choice depends on the target model.
Does RX 7900 XT work with Ollama on Windows?+
Yes: the Ollama documentation lists it for Windows, with a ROCm v7 or HIP7 stack, and for Linux with the ROCm v7 driver. AMD lists it in ROCm, targeting gfx1100, on Ubuntu 24.04.4, 22.04.5, RHEL 10.1, and RHEL 9.7. Vulkan remains available as a fallback.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.