Intermediate 11 minRadeon RX 7000

Which LLM on Radeon RX 7800 XT (16 GB) ?

Direct response

The Radeon RX 7800 XT is a good choice for local AI mainly because of its 16 GB of VRAM and 624 GB/s of bandwidth: it loads 20- to 24-billion-parameter models in Q4, which 12 GB cards cannot do. A public benchmark on Llama 2 7B in Q4_0 reports 118 tokens per second for generation under Vulkan. Its price matters less than the cost per GB of VRAM, which you can compare on /prix-gpu-ia.

A good graphics card deal is judged by what the desired model fits on it, not by the listed price. RX 7800 XT is one of the few 16 GB cards from the RDNA 3 generation officially supported by ROCm and Ollama. This guide shows what it can run, how fast according to published measurements, and how to evaluate an offer without relying on a price that changes every week.

Choosing a machine? Our picks by budget →

By Mohamed Meguedmi·Update 2026-09-30·Tested on Windows, macOS, and Linux
Recommended hardware

Buying alternative for this guide: Radeon RX 9070 XT 16GB (ASUS Prime OC).

Why this choice? Our complete guide on Radeon RX 9070 XT 16GB (ASUS Prime OC) →

Compare all options by budget, from €800 to €3,500 →

On the go: which laptop for local AI →

Affiliate links — possible commission at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.

#Why the RX 7800 XT is a good choice for a local LLM

The RX 7800 XT combines three rare strengths at this price point. Its 16 GB of GDDR6 on a 256-bit bus gives it 624 GB/s of bandwidth, the factor that limits an LLM's generation speed. AMD officially lists it in ROCm (RDNA 3, gfx1101 target), and Ollama cites it for both Linux and Windows. Finally, public measurements place it at 118.27 tokens per second when generating on a 7B model in Q4_0 under Vulkan—faster than a RX 7900 GRE in the same series of measurements. Its drawback compared with NVIDIA is not generation speed but prompt processing, which is three times slower than a RTX 4070 Ti Super.

Architecture
RDNA 3, Navi 32 in chiplets: one GCD and four MCDs, launched on September 6, 2023. It is therefore not monolithic.
Compute
3,840 stream processors, 60 compute units.
Memory
16 GB GDDR6, 256-bit bus, 624 GB/s.
Power consumption
263 W of board power.
Pricing
Not set here: see /prix-gpu-ia, which records the lowest price every Monday and Thursday.

#Evaluating an offer without a fixed price

The term “good deal” implies a price, and a card’s price changes from week to week. Rather than cite one that will be wrong in ten days, here’s the method. Calculate the cost per GB of VRAM: the price divided by 16 for the 7800 XT, or by 12 for a 12 GB card. The site’s tracker publishes this figure for each card. Then ask yourself the only question that matters: does the extra capacity unlock a model you actually want to use? Moving from 12 to 16 GB unlocks 20B to 24B models; moving from 16 to 24 GB unlocks 30B and larger models.

Benchmarks for evaluating an offering (calculation example, not prices)
SituationCalculation to makeConclusion
Deal on a 7800 XTPrice ÷ 16 GBCompare the cost per GB of 12 GB and 24 GB cards on /prix-gpu-ia
Price difference compared with a 12 GB cardGap ÷ 4 GB additionalBelow the average cost per GB, the 16 GB model is a great deal
Difference from a 16 GB RTXWhat CUDA is worth for youPay for it only if a tool requires it or prompt reading matters
Used cardTest with a 12 to 14 GB model before buyingReject it without a real load test

An example of the reasoning, with no pricing implication: if the gap between a 12 GB card and the 7800 XT amounts to what you would have spent anyway on a 24B model, the question is settled. Conversely, if you only use 8B to 12B models, those 4 GB serve no purpose, and the RX 7700 XT or the RX 6700 XT are sufficient.

#What fits in 16 GB of VRAM

With 16 GB, allow about 15 GB for the model and its KV cache when the card isn’t displaying the screen. The QuelLLM catalog lists the Q4 weights below; the context is added on top, and headroom quickly becomes the limiting factor on large models.

What a 7800 XT can handle (weights only)
ModelWeightsHeadroom for contextVerdict
Gemma 4 12B in Q8≈ 13 GB≈ 2 GBFits, short context
gpt-oss 20B in Q4≈ 13 GB≈ 2 GBHolds up, short- to medium-length context
Devstral Small 2 24B in Q4≈ 14 GB≈ 1 GBToo cramped, very short context
Gemma 4 26B-A4B in Q4≈ 16 GBnoneToo close to call
Granite 4.1 30B in Q4≈ 17 GBnoneDoesn't fit

The key word in the table is headroom: a 24B in Q4 fits, but only with a context of a few thousand tokens. For coding with long files, KV cache quantization frees up one to two gigabytes, and a 20 GB card like the RX 7900 XT solves the problem. The site's table on 16 GB models, independent of the card brand, details the quantization trade-offs.

A word about mixture-of-experts (MoE) models such as gpt-oss 20B: only some of the parameters work on each token, which speeds up generation, but all weights must reside in memory. It is the total size, not the active size, that determines whether the model fits in 16 GB. That explains why a 20B MoE appears alongside a 12B dense model in the table above, with the same weight of about 13 GB.

#Measured throughput: Vulkan or ROCm

The figures below are not measurements from the site: they come from two public discussions in the llama.cpp repository, one for Vulkan and the other for ROCm, both testing Llama 2 7B in Q4_0 with llama-bench. The llama.cpp commits, systems, and drivers differ between runs, so the gap between the two series is an indication, not a definitive ranking.

RX 7800 XT on Llama 2 7B Q4_0 (public llama.cpp measurements)
BackendFlash Attentionpp512 processingGeneration tg128
Vulkanno2 017,33 t/s118,27 t/s
Vulkanyes2 197,05 t/s124,86 t/s
ROCmno2 151,81 t/s100,94 t/s
ROCmyes2 304,63 t/s95,99 t/s

Two takeaways stand out. First, in these tests, Vulkan generates faster than ROCm on this card (118 versus 101 tokens per second without Flash Attention), while ROCm reads the prompt a little faster. Second, Flash Attention speeds up Vulkan but slightly slows generation under ROCm. If your priority is conversational response speed, try both backends with your model rather than assuming ROCm wins.

Performance relative to the theoretical ceiling helps put things in perspective: 624 GB/s divided by 3.82 GB of weights gives 163 tokens per second; the card reaches 72% of that under Vulkan. Applied to larger models, this ratio provides rough estimates to confirm on your own system.

Estimated at 72% of the 624 GB/s ceiling (calculation, not measurement)
Model (Q4)WeightsTheoretical ceilingOrder of magnitude
Llama 3.1 8B≈ 6 GB≈ 104 tokens/s≈ 75 tokens/s
Gemma 4 12B≈ 7 GB≈ 89 tokens/s≈ 64 tokens/s
Phi-4 14B≈ 9 GB≈ 69 tokens/s≈ 50 tokens/s
Devstral Small 2 24B≈ 14 GB≈ 45 tokens/s≈ 32 tokens/s

#Against NVIDIA cards with 16 GB: prompt processing settles the question

The comparison with 16 GB RTX cards comes down to two metrics. For generation, the RX 7800 XT remains in the same range as the NVIDIA cards measured under Vulkan. For prompt processing, the gap is very large: three times slower than a RTX 4070 Ti Super. This second figure matters for use cases involving long documents or long conversation histories, where you're waiting for the first word. It matters little for a short chat or a coding assistant with a modest context.

RX 7800 XT and 16 GB RTX under Vulkan (Llama 2 7B Q4_0, without Flash Attention)
CardVRAMpp512 processingGeneration tg128
RX 7800 XT16 GB2 017,33 t/s118,27 t/s
RTX 4070 Ti Super16 GB6 099,18 t/s129,45 t/s
RTX 5070 Ti16 GB6 213,63 t/s135,63 t/s

#How to fit a 24B model into 16 GB: useful levers

A 24-billion-parameter model in Q4 weighs about 14 GB, leaving barely a gigabyte for the KV cache and system. Four levers prevent overflow onto the processor, which would sharply reduce speed. The first is to limit the context window to what you actually need, because the cache grows with every token. The second is to quantize the KV cache, freeing up one to two gigabytes. The third is to close applications that reserve VRAM, such as a browser with hardware acceleration or a game. The fourth is to connect the display to the processor’s integrated graphics when available, returning to the AMD card the memory used by the display.

Context
Set the window to match your actual needs: 4,000 to 8,000 tokens are enough for most conversations.
KV cache
Quantize it to 8-bit to free up VRAM, while monitoring quality on your tasks.
Display
Use the integrated graphics card’s video output when available.
Verification
After loading, ollama ps should show 100% GPU; any other proportion indicates overflow.

The guide to the context window explains why a long history consumes as much memory as the model itself on smaller cards.

#Getting started

  1. 01
    Choose the system
    On Linux, AMD supports Radeon RX 7000 only on Ubuntu 24.04.4, Ubuntu 22.04.5, RHEL 10.1, and RHEL 9.7. On Windows, Ollama relies on a ROCm v7 or HIP7 stack.
  2. 02
    Install Ollama
    Follow the Ollama documentation: ROCm v7 is required on Linux. Vulkan is enabled by default and serves as a fallback.
  3. 03
    Load a model that fits
    Try a 12B in Q8 or a 20B in Q4 before attempting a 24B. Check with ollama ps that it is entirely on the GPU.
  4. 04
    Compare Vulkan and ROCm
    Run llama-bench with both backends on your model, with and without Flash Attention, and keep the faster one for your use case.
Terminal
ollama ps
./llama-bench -m llama-2-7b.Q4_0.gguf -ngl 100 -fa 0,1

#2026 verdict

A good option if
You’re targeting 20B to 24B models, and the VRAM cost per GB on /prix-gpu-ia is lower than on nearby 12 GB cards.
Not a good choice if
You’re sticking with 8B to 12B models: a 12 GB card is enough, and you’re paying for unused gigabytes.
Avoid
Buying a 12 to 14 GB model without testing whether it loads when the card is used.

#Frequently asked questions

FAQ
RX 7800 XT: is it a good deal for running LLMs?+
Yes, if you are targeting models with 20 to 24 billion parameters: its 16 GB accepts them in Q4, whereas 12 GB cards fail. Evaluate the offering based on the cost per GB of VRAM, published on the site's pricing page, rather than the listed price alone.
Is the RX 7800 XT good for LLMs in 2026?+
It generates 118 tokens per second on a 7B model in Q4_0 under Vulkan, according to a public benchmark, and its 16 GB can load a 24B model in Q4. Its weakness compared with NVIDIA is prompt processing, which is about three times slower than a RTX 4070 Ti Super in these measurements.
Vulkan or ROCm on RX 7800 XT?+
In the two public discussions in the llama.cpp repository, Vulkan generates faster (118 versus 101 tokens per second without Flash Attention), while ROCm reads the prompt slightly faster. Configurations differ, so test both with your model and keep the faster one.
RX 7800 XT or RX 7700 XT for local AI?+
The 7800 XT adds 4 GB of VRAM, providing access to 20B to 24B models, and 624 GB/s of bandwidth versus 432. If you stick to 8B to 14B models, the 7700 XT is enough; otherwise, the price difference is generally justified.
Which models fit on a RX 7800 XT with 16 GB?+
Gemma 4 12B in Q8 and gpt-oss 20B in Q4 weigh about 13 GB, Devstral Small 2 24B in Q4 about 14 GB, with a short context. A 26B to 30B model in Q4 exceeds 16 GB of weights and does not fit entirely.
Is the RX 7800 XT supported by Ollama?+
Yes, on both Linux and Windows according to the Ollama documentation, with the ROCm v7 driver on Linux. AMD also lists it in ROCm, target gfx1101, but only on Ubuntu 24.04.4, Ubuntu 22.04.5, RHEL 10.1, or RHEL 9.7. Vulkan, enabled by default in Ollama, serves as a fallback.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.