Which LLM on Radeon RX 7900 XT (20 GB) ?
The Radeon RX 7900 XT (20 GB GDDR6, 800 GB/s, 315 W) loads 24B to 26B models in Q4 without spilling, with a comfortable context, and can still handle a 30B model at the cost of a short context. It's memory, not speed, that sets it apart: on Llama 2 7B, public benchmarks report 123 tokens per second under Vulkan, barely faster than a RX 7800 XT with 16 GB, and far behind the RX 7900 XTX.
Twenty gigabytes of VRAM is rare capacity on a consumer card, and it changes what you can run: it’s the difference between a cramped 24B and a 24B with context. This guide separates what those 20 GB actually provide from what 800 GB/s of bandwidth might suggest, using cited public measurements and correcting misconceptions about this card.
Choosing a machine? Our picks by budget →
Buying alternative for this guide: Radeon RX 9070 XT 16GB (ASUS Prime OC).
Why this choice? Our complete guide on Radeon RX 9070 XT 16GB (ASUS Prime OC) →
Compare all options by budget, from €800 to €3,500 →
On the go: which laptop for local AI →
Affiliate links — possible commission at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.
#What a RX 7900 XT does in 2026
The RX 7900 XT is for anyone who wants 24- to 26-billion-parameter models without constraining the context. Its 20 GB accommodates Devstral Small 2 24B (about 14 GB in Q4) with 6 GB to spare, versus 1 GB on a 16 GB card, and the MoE Gemma 4 26B-A4B (about 16 GB) with 4 GB to spare. In speed, however, it disappoints anyone counting on its 800 GB/s: 123.18 tokens per second on the reference 7B under Vulkan, versus 118.27 for a RX 7800 XT whose bandwidth is 22% lower. AMD officially lists it in ROCm, and Ollama supports it on Linux and Windows.
- Architecture
- RDNA 3, Navi 31 with one GCD and five MCDs, launched on December 13, 2022.
- Compute
- 5,376 stream processors, 84 compute units.
- Memory
- 20 GB GDDR6, 320-bit bus, 800 GB/s.
- Power consumption
- 315 W of board power: check the power supply recommended by your model’s manufacturer.
- Pricing
- Not specified here: see /prix-gpu-ia.
#The 20 GB memory budget
Allow about 19 GB for the model and its KV cache when the card is not driving the display. The table subtracts the Q4 size from the QuelLLM catalog from this budget to show the headroom available for context. This headroom is the real criterion: a model that “fits” with 1 GB left has room for only a few thousand tokens of context.
| Model | Weights | Headroom for context | Verdict |
|---|---|---|---|
| gpt-oss 20B | ≈ 13 GB | ≈ 6 GB | Very comfortable, long context |
| Devstral Small 2 24B | ≈ 14 GB | ≈ 5 GB | Comfortable, long context possible |
| Gemma 4 26B-A4B (MoE) | ≈ 16 GB | ≈ 3 GB | Fits, medium context |
| Granite 4.1 30B | ≈ 17 GB | ≈ 2 GB | Fits, short context |
| Gemma 4 31B | ≈ 18 GB | ≈ 1 GB | Cramped, very short context |
Mixture-of-experts models deserve a note. Only a fraction of an MoE's parameters work on each token, making generation faster than with a dense model of the same size, but all the weights still have to fit in memory. The 26B-A4B is the kind of model where 20 GB makes the difference: too large for 16 GB with context, comfortable in 20. For even larger models, such as the Qwen3.6 35B-A3B family, consult the dedicated guide rather than guessing: the exact size depends on the quantization used.
To convert the margin into a number of context tokens, you need the KV-cache cost per token, which depends on the model architecture: number of layers, attention heads, and cache precision. So there is no universal tokens-per-gigabyte rule. The reliable method is measurement: load the model with a modest context, note the VRAM usage with rocm-smi, increase the window, and repeat. The site's VRAM calculator automates this estimate for models in the catalog, and quantizing the KV cache to 8-bit roughly halves the context's share.
- Qwen3.6 35B-A3B locally: testing and VRAM requirements
- Which LLM for 24 GB of VRAM, for models larger than 20 GB
- Quantize the KV cache to save VRAM
#Measured throughput: bandwidth is not everything
The figures come from public discussions in the llama.cpp repository, which use Llama 2 7B in Q4_0 and llama-bench; they are not measurements from the site. Under Vulkan, the RX 7900 XT reads a 512-token prompt at 2,941.58 tokens per second and generates at 123.18. With Flash Attention, the figures are 2,701.13 and 120.62. Under ROCm, the discussion reports 3,098.38 for reading and 116.15 for generation. The two sets converge: generation tops out at around 120 tokens per second.
The theoretical bandwidth ceiling, 800 GB/s divided by 3.82 GB of weights, is about 209 tokens per second; the card reaches only 59% under Vulkan. This is the lowest efficiency among the AMD cards in this series of measurements (83% for a RX 6700 XT, 77% for a RX 7900 GRE, 72% for a RX 7800 XT). The likely explanation is that a 7-billion-parameter model is too small to saturate the memory of a large chip; better efficiency can be expected on larger models, but none of the measurements in these series proves it, and it would be imprudent to announce throughput figures for a 24B.
| Backend | Flash Attention | pp512 processing | Generation tg128 |
|---|---|---|---|
| Vulkan | no | 2 941,58 t/s | 123,18 t/s |
| Vulkan | yes | 2 701,13 t/s | 120,62 t/s |
| ROCm | no | 3 098,38 t/s | 116,15 t/s |
For a ballpark estimate on large models, a cautious bound is to apply the measured 59% efficiency to the theoretical ceiling: a 24B model in Q4 at 14 GB gives 800 divided by 14, or about 57 tokens per second at the ceiling, and therefore roughly 34 tokens per second. This is a calculation, not a measurement; if efficiency is better on a large model, the actual throughput will be higher.
#Facing RX 7900 XTX: the real gap is much larger than people say
You often read that the 7900 XT is “10 to 12% slower” than the 7900 XTX. Public benchmarks tell a different story. On the reference 7B model under Vulkan, the XTX generates 182.63 tokens per second, versus 123.18 for the XT: 48% more. The stream processor ratio (6,144 versus 5,376) isn’t enough to explain it; memory bandwidth and the chip itself create the gap. What the XT preserves is the entry price: with 20 GB, it covers most of what 24B to 26B models require.
| Metric | RX 7900 XT | RX 7900 XTX |
|---|---|---|
| VRAM | 20 GB | 24 GB |
| Vulkan, tg128 generation | 123,18 t/s | 182,63 t/s |
| Vulkan, pp512 prompt processing | 2 941,58 t/s | 3 726,99 t/s |
This correction changes the trade-off: if generation speed matters to you, the 48% gap outweighs the additional 4 GB of VRAM, whereas if you mainly want capacity for context, the XT delivers most of what the XTX offers for a 24B model. On a 30B model or larger, however, the XTX's 24 GB eliminates context compromises, and that's where its price premium is justified.
#Versus the RTX 4070 Ti Super: more memory, less speed
The RTX 4070 Ti Super offers 16 GB, which is 4 less. On the reference 7B model under Vulkan, it generates at 129.45 tokens per second, 5% better than the RX 7900 XT, and processes the prompt at 6,099 tokens per second, twice as fast. So the XT is not “faster,” contrary to what you sometimes read: it is larger. Its advantage is capacity, which enables 24B to 26B models with context.
| Criterion | RX 7900 XT | RTX 4070 Ti Super |
|---|---|---|
| VRAM | 20 GB | 16 GB |
| pp512 processing | 2 941,58 t/s | 6 099,18 t/s |
| Generation tg128 | 123,18 t/s | 129,45 t/s |
#When a model exceeds 20 GB
A model that’s too large doesn’t cause an error: Ollama and llama.cpp place as many layers as possible on the card and leave the rest to the CPU. In llama.cpp, the -ngl option sets the number of layers sent to the GPU, and the value 100 used in the measurements above means “all.” Reducing this value frees up VRAM, at the cost of slower generation, since each token must then pass through system memory, which is much slower than GDDR6. The rule of thumb: as long as ollama ps shows 100% GPU, you’re in the fast zone; as soon as a proportion appears on the CPU side, speed drops in steps.
#20 GB, 16 GB, or 24 GB: how to choose
| Intended use | 16 GB | 20 GB (RX 7900 XT) | 24 GB |
|---|---|---|---|
| Chat with an 8B to 14B model | Sufficient | Unnecessary | Unnecessary |
| 24B coding assistant, long context | Cramped | Comfortable | Comfortable |
| 26B MoE with medium context | Too close to call | Sufficient | Sufficient |
| 30B to 35B model with long context | Impossible | Just | Required |
The principle: pay for the capacity your largest model requires, not for what you might use someday. If your target is a 24B model with context, 20 GB is the first tier where it becomes comfortable; beyond that, each additional gigabyte must be justified by a specific model.
#Getting started
- 01Choose the systemAMD supports Radeon RX 7000 only on Ubuntu 24.04.4, Ubuntu 22.04.5, RHEL 10.1, and RHEL 9.7. On Windows, Ollama uses a ROCm v7 or HIP7 stack; Vulkan serves as a fallback.
- 02Install the driver and OllamaOn Linux, Ollama requires the ROCm v7 driver. Install it with the amdgpu-install utility, then Ollama.
- 03Load a 24B modelChoose a Devstral Small 2 24B in Q4_K_M, run it, and check with ollama ps that it is using the GPU at 100%.
- 04Adjust the contextIncrease the context window in stages while monitoring VRAM: the goal is to stay below 19 GB total.
- 05Compare Vulkan and ROCmMeasure with llama-bench on your model. The two backends are close on this card in publicly available benchmarks.
#2026 verdict
- A good choice
- If you want to run 24B to 26B models with context and the cost per GB of VRAM on /prix-gpu-ia is favorable.
- A middle-ground choice
- If speed comes first: generation tops out near the level of a 16 GB card, and the RTX 4070 Ti Super reads prompts twice as fast.
- An insufficient choice
- For 30B to 35B models with a long context, where 24 GB becomes necessary.
- Local AI graphics card prices, recorded twice a week
- Ollama documentation: supported AMD cards
- ROCm GPU compatibility, AMD documentation
- Vulkan scoreboard for the llama.cpp repository
#Frequently asked questions
Is the RX 7900 XT good for local LLMs?+
RX 7900 XT or RX 7900 XTX for local AI?+
Is 20 GB of VRAM enough for a 30B model?+
How many tokens per second does a RX 7900 XT produce?+
RX 7900 XT or RTX 4070 Ti Great for an LLM?+
Does RX 7900 XT work with Ollama on Windows?+
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.