Which LLM on Radeon RX 7800 XT (16 GB) ?
The Radeon RX 7800 XT is a good choice for local AI mainly because of its 16 GB of VRAM and 624 GB/s of bandwidth: it loads 20- to 24-billion-parameter models in Q4, which 12 GB cards cannot do. A public benchmark on Llama 2 7B in Q4_0 reports 118 tokens per second for generation under Vulkan. Its price matters less than the cost per GB of VRAM, which you can compare on /prix-gpu-ia.
A good graphics card deal is judged by what the desired model fits on it, not by the listed price. RX 7800 XT is one of the few 16 GB cards from the RDNA 3 generation officially supported by ROCm and Ollama. This guide shows what it can run, how fast according to published measurements, and how to evaluate an offer without relying on a price that changes every week.
Choosing a machine? Our picks by budget →
Buying alternative for this guide: Radeon RX 9070 XT 16GB (ASUS Prime OC).
Why this choice? Our complete guide on Radeon RX 9070 XT 16GB (ASUS Prime OC) →
Compare all options by budget, from €800 to €3,500 →
On the go: which laptop for local AI →
Affiliate links — possible commission at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.
#Why the RX 7800 XT is a good choice for a local LLM
The RX 7800 XT combines three rare strengths at this price point. Its 16 GB of GDDR6 on a 256-bit bus gives it 624 GB/s of bandwidth, the factor that limits an LLM's generation speed. AMD officially lists it in ROCm (RDNA 3, gfx1101 target), and Ollama cites it for both Linux and Windows. Finally, public measurements place it at 118.27 tokens per second when generating on a 7B model in Q4_0 under Vulkan—faster than a RX 7900 GRE in the same series of measurements. Its drawback compared with NVIDIA is not generation speed but prompt processing, which is three times slower than a RTX 4070 Ti Super.
- Architecture
- RDNA 3, Navi 32 in chiplets: one GCD and four MCDs, launched on September 6, 2023. It is therefore not monolithic.
- Compute
- 3,840 stream processors, 60 compute units.
- Memory
- 16 GB GDDR6, 256-bit bus, 624 GB/s.
- Power consumption
- 263 W of board power.
- Pricing
- Not set here: see /prix-gpu-ia, which records the lowest price every Monday and Thursday.
#Evaluating an offer without a fixed price
The term “good deal” implies a price, and a card’s price changes from week to week. Rather than cite one that will be wrong in ten days, here’s the method. Calculate the cost per GB of VRAM: the price divided by 16 for the 7800 XT, or by 12 for a 12 GB card. The site’s tracker publishes this figure for each card. Then ask yourself the only question that matters: does the extra capacity unlock a model you actually want to use? Moving from 12 to 16 GB unlocks 20B to 24B models; moving from 16 to 24 GB unlocks 30B and larger models.
| Situation | Calculation to make | Conclusion |
|---|---|---|
| Deal on a 7800 XT | Price ÷ 16 GB | Compare the cost per GB of 12 GB and 24 GB cards on /prix-gpu-ia |
| Price difference compared with a 12 GB card | Gap ÷ 4 GB additional | Below the average cost per GB, the 16 GB model is a great deal |
| Difference from a 16 GB RTX | What CUDA is worth for you | Pay for it only if a tool requires it or prompt reading matters |
| Used card | Test with a 12 to 14 GB model before buying | Reject it without a real load test |
An example of the reasoning, with no pricing implication: if the gap between a 12 GB card and the 7800 XT amounts to what you would have spent anyway on a 24B model, the question is settled. Conversely, if you only use 8B to 12B models, those 4 GB serve no purpose, and the RX 7700 XT or the RX 6700 XT are sufficient.
#What fits in 16 GB of VRAM
With 16 GB, allow about 15 GB for the model and its KV cache when the card isn’t displaying the screen. The QuelLLM catalog lists the Q4 weights below; the context is added on top, and headroom quickly becomes the limiting factor on large models.
| Model | Weights | Headroom for context | Verdict |
|---|---|---|---|
| Gemma 4 12B in Q8 | ≈ 13 GB | ≈ 2 GB | Fits, short context |
| gpt-oss 20B in Q4 | ≈ 13 GB | ≈ 2 GB | Holds up, short- to medium-length context |
| Devstral Small 2 24B in Q4 | ≈ 14 GB | ≈ 1 GB | Too cramped, very short context |
| Gemma 4 26B-A4B in Q4 | ≈ 16 GB | none | Too close to call |
| Granite 4.1 30B in Q4 | ≈ 17 GB | none | Doesn't fit |
The key word in the table is headroom: a 24B in Q4 fits, but only with a context of a few thousand tokens. For coding with long files, KV cache quantization frees up one to two gigabytes, and a 20 GB card like the RX 7900 XT solves the problem. The site's table on 16 GB models, independent of the card brand, details the quantization trade-offs.
A word about mixture-of-experts (MoE) models such as gpt-oss 20B: only some of the parameters work on each token, which speeds up generation, but all weights must reside in memory. It is the total size, not the active size, that determines whether the model fits in 16 GB. That explains why a 20B MoE appears alongside a 12B dense model in the table above, with the same weight of about 13 GB.
#Measured throughput: Vulkan or ROCm
The figures below are not measurements from the site: they come from two public discussions in the llama.cpp repository, one for Vulkan and the other for ROCm, both testing Llama 2 7B in Q4_0 with llama-bench. The llama.cpp commits, systems, and drivers differ between runs, so the gap between the two series is an indication, not a definitive ranking.
| Backend | Flash Attention | pp512 processing | Generation tg128 |
|---|---|---|---|
| Vulkan | no | 2 017,33 t/s | 118,27 t/s |
| Vulkan | yes | 2 197,05 t/s | 124,86 t/s |
| ROCm | no | 2 151,81 t/s | 100,94 t/s |
| ROCm | yes | 2 304,63 t/s | 95,99 t/s |
Two takeaways stand out. First, in these tests, Vulkan generates faster than ROCm on this card (118 versus 101 tokens per second without Flash Attention), while ROCm reads the prompt a little faster. Second, Flash Attention speeds up Vulkan but slightly slows generation under ROCm. If your priority is conversational response speed, try both backends with your model rather than assuming ROCm wins.
Performance relative to the theoretical ceiling helps put things in perspective: 624 GB/s divided by 3.82 GB of weights gives 163 tokens per second; the card reaches 72% of that under Vulkan. Applied to larger models, this ratio provides rough estimates to confirm on your own system.
| Model (Q4) | Weights | Theoretical ceiling | Order of magnitude |
|---|---|---|---|
| Llama 3.1 8B | ≈ 6 GB | ≈ 104 tokens/s | ≈ 75 tokens/s |
| Gemma 4 12B | ≈ 7 GB | ≈ 89 tokens/s | ≈ 64 tokens/s |
| Phi-4 14B | ≈ 9 GB | ≈ 69 tokens/s | ≈ 50 tokens/s |
| Devstral Small 2 24B | ≈ 14 GB | ≈ 45 tokens/s | ≈ 32 tokens/s |
#Against NVIDIA cards with 16 GB: prompt processing settles the question
The comparison with 16 GB RTX cards comes down to two metrics. For generation, the RX 7800 XT remains in the same range as the NVIDIA cards measured under Vulkan. For prompt processing, the gap is very large: three times slower than a RTX 4070 Ti Super. This second figure matters for use cases involving long documents or long conversation histories, where you're waiting for the first word. It matters little for a short chat or a coding assistant with a modest context.
| Card | VRAM | pp512 processing | Generation tg128 |
|---|---|---|---|
| RX 7800 XT | 16 GB | 2 017,33 t/s | 118,27 t/s |
| RTX 4070 Ti Super | 16 GB | 6 099,18 t/s | 129,45 t/s |
| RTX 5070 Ti | 16 GB | 6 213,63 t/s | 135,63 t/s |
- RTX 4070 Ti Super, the 16 GB alternative NVIDIA
- Radeon RX 7900 GRE, another 16 GB card from the same generation
#How to fit a 24B model into 16 GB: useful levers
A 24-billion-parameter model in Q4 weighs about 14 GB, leaving barely a gigabyte for the KV cache and system. Four levers prevent overflow onto the processor, which would sharply reduce speed. The first is to limit the context window to what you actually need, because the cache grows with every token. The second is to quantize the KV cache, freeing up one to two gigabytes. The third is to close applications that reserve VRAM, such as a browser with hardware acceleration or a game. The fourth is to connect the display to the processor’s integrated graphics when available, returning to the AMD card the memory used by the display.
- Context
- Set the window to match your actual needs: 4,000 to 8,000 tokens are enough for most conversations.
- KV cache
- Quantize it to 8-bit to free up VRAM, while monitoring quality on your tasks.
- Display
- Use the integrated graphics card’s video output when available.
- Verification
- After loading, ollama ps should show 100% GPU; any other proportion indicates overflow.
The guide to the context window explains why a long history consumes as much memory as the model itself on smaller cards.
#Getting started
- 01Choose the systemOn Linux, AMD supports Radeon RX 7000 only on Ubuntu 24.04.4, Ubuntu 22.04.5, RHEL 10.1, and RHEL 9.7. On Windows, Ollama relies on a ROCm v7 or HIP7 stack.
- 02Install OllamaFollow the Ollama documentation: ROCm v7 is required on Linux. Vulkan is enabled by default and serves as a fallback.
- 03Load a model that fitsTry a 12B in Q8 or a 20B in Q4 before attempting a 24B. Check with ollama ps that it is entirely on the GPU.
- 04Compare Vulkan and ROCmRun llama-bench with both backends on your model, with and without Flash Attention, and keep the faster one for your use case.
#2026 verdict
- A good option if
- You’re targeting 20B to 24B models, and the VRAM cost per GB on /prix-gpu-ia is lower than on nearby 12 GB cards.
- Not a good choice if
- You’re sticking with 8B to 12B models: a 12 GB card is enough, and you’re paying for unused gigabytes.
- Avoid
- Buying a 12 to 14 GB model without testing whether it loads when the card is used.
- Ollama documentation: supported AMD cards
- ROCm GPU compatibility, AMD documentation
- Vulkan scoreboard for the llama.cpp repository
- ROCm scoreboard for the llama.cpp repository
#Frequently asked questions
RX 7800 XT: is it a good deal for running LLMs?+
Is the RX 7800 XT good for LLMs in 2026?+
Vulkan or ROCm on RX 7800 XT?+
RX 7800 XT or RX 7700 XT for local AI?+
Which models fit on a RX 7800 XT with 16 GB?+
Is the RX 7800 XT supported by Ollama?+
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.