Which LLM on Radeon RX 7900 GRE (16 GB) ?
The Radeon RX 7900 GRE (16 GB GDDR6, 576 GB/s, 260 W) runs 20- to 24-billion-parameter models in Q4, like the RX 7800 XT, but without gaining generation speed: public measurements on Llama 2 7B report 116 tokens per second versus 118 for the 7800 XT under Vulkan. Its advantage lies elsewhere, in prompt processing under Vulkan, which is slightly faster (2,336 versus 2,017 tokens per second).
The 7900 GRE looks like a 7800 XT with stronger compute performance but less memory, and the benchmarks show it. This guide explains what its 16 GB can handle, where it differs from its two direct neighbors, the RX 7800 XT and RX 9070, and what use case makes it worth the price.
Choosing a machine? Our picks by budget →
Buying alternative for this guide: Radeon RX 9070 XT 16GB (ASUS Prime OC).
Why this choice? Our complete guide on Radeon RX 9070 XT 16GB (ASUS Prime OC) →
Compare all options by budget, from €800 to €3,500 →
On the go: which laptop for local AI →
Affiliate links — commission possible at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.
#What a RX 7900 GRE does in 2026
The RX 7900 GRE handles anything under 15 GB in Q4: Gemma 4 12B through Q8, gpt-oss 20B, and Devstral Small 2 24B with a short context. Its generation speed, capped by its 576 GB/s, is around 116 tokens per second on the benchmark scoreboards’ reference 7B model under Vulkan. That is the key takeaway: its 5,120 stream processors, versus 3,840 for the 7800 XT, speed up prompt processing but not generation, because it has neither more memory nor more bandwidth than its counterpart. AMD officially lists it in ROCm, and Ollama supports it on Linux and Windows.
- Architecture
- RDNA 3, Navi 31 (the RX 7900 XTX chip, reduced here), 80 compute units and 5,120 stream processors.
- Memory
- 16 GB GDDR6 with a 256-bit bus, providing 576 GB/s.
- Power consumption
- 260 W of board power.
- Availability
- Initially reserved for China (July 27, 2023), marketed internationally on February 27, 2024, at a launch price of $549.
- Current price
- See /prix-gpu-ia: the price changes every week.
#7900 GRE vs. 7800 XT: memory decides
Both cards offer 16 GB. The 7800 XT has 624 GB/s of bandwidth, while the 7900 GRE has 576 GB/s, or 8% less. During generation, when every token rereads all the weights, bandwidth is what matters: the 7800 XT therefore slightly outperforms the GRE in the llama.cpp Vulkan discussion, and more clearly with Flash Attention enabled. The GRE regains the advantage for prompt processing under Vulkan (16% more), where compute matters. Under ROCm, both public series show the 7800 XT ahead on both measurements.
| Metric | RX 7900 GRE | RX 7800 XT |
|---|---|---|
| Vulkan, pp512 processing, without FA | 2 336,31 t/s | 2 017,33 t/s |
| Vulkan, tg128 generation, without FA | 116,11 t/s | 118,27 t/s |
| Vulkan, tg128 generation, with FA | 110,50 t/s | 124,86 t/s |
| ROCm, pp512 processing, without FA | 1 456,98 t/s | 2 151,81 t/s |
| ROCm, tg128 generation, without FA | 96,07 t/s | 100,94 t/s |
These results come from different contributors, with different llama.cpp commits: the gap is an order of magnitude, not a laboratory comparison. They nevertheless agree that the GRE does not generate faster than a 7800 XT. If you find both at the same price, choose the 7800 XT for conversational use, and the GRE for use with long prompts.
#What fits in 16 GB of VRAM
Allow about 15 GB for the model and its KV cache when the card is not powering the display. The Q4 weights come from the QuelLLM catalog; the context is added on top.
| Model | Weights | Headroom for context | Verdict |
|---|---|---|---|
| Gemma 4 12B in Q4 | ≈ 7 GB | ≈ 8 GB | Very comfortable, long context |
| Gemma 4 12B in Q8 | ≈ 13 GB | ≈ 2 GB | Fits, short context |
| gpt-oss 20B in Q4 | ≈ 13 GB | ≈ 2 GB | Holds up, short- to medium-length context |
| Devstral Small 2 24B in Q4 | ≈ 14 GB | ≈ 1 GB | Too cramped |
| Gemma 4 31B in Q4 | ≈ 18 GB | none | Doesn't fit; you need 20 to 24 GB |
The 16 GB boundary therefore lies between the 24B, which fits, and the 31B, which does not. To cross this boundary without moving to a 24 GB card, the RX 7900 XT at 20 GB provides an intermediate margin; the dedicated guide explains what it offers. For all 16 GB cards, regardless of brand, the VRAM-by-VRAM page compares the models.
#Measured throughput and bandwidth ceiling
None of the figures on this page are measurements of the site. The values come from public discussions in the llama.cpp repository, all of which use Llama 2 7B in Q4_0 (3.56 GiB) and llama-bench. Prompt processing (pp512) measures the ingestion speed for 512 tokens; generation (tg128) measures writing speed. For the RX 7900 GRE under Vulkan, the discussion reports 116.11 tokens per second during generation without Flash Attention and 110.50 with it. Under ROCm, it reports 96.07.
The theoretical generation ceiling, bandwidth divided by weight size, is 576 GB/s here for 3.82 GB, or about 151 tokens per second. The card reaches 77% under Vulkan without Flash Attention. Applying this efficiency to larger models gives a ballpark figure: it is not a measurement, and another driver or another version of llama.cpp will change it.
| Model (Q4) | Weights | Theoretical ceiling | Order of magnitude |
|---|---|---|---|
| Llama 3.1 8B | ≈ 6 GB | ≈ 96 tokens/s | ≈ 72 tokens/s |
| Gemma 4 12B | ≈ 7 GB | ≈ 82 tokens/s | ≈ 62 tokens/s |
| Phi-4 14B | ≈ 9 GB | ≈ 64 tokens/s | ≈ 48 tokens/s |
| Devstral Small 2 24B | ≈ 14 GB | ≈ 41 tokens/s | ≈ 31 tokens/s |
#What reading the prompt changes in practice
Prompt processing seems abstract until you convert it into seconds of waiting. The calculation is simple: prompt token count divided by pp512 speed. It assumes that the speed measured on 512 tokens remains constant for a longer prompt, which is only approximately true, because processing slows as the context grows. The times below are therefore rough estimates for a document of about 4,000 tokens, or roughly ten pages.
| Card and backend | pp512 speed | Time to first token |
|---|---|---|
| RX 7900 GRE, Vulkan | 2 336 t/s | ≈ 1,7 s |
| RX 7800 XT, Vulkan | 2 017 t/s | ≈ 2,0 s |
| RX 7900 GRE, ROCm | 1 457 t/s | ≈ 2,7 s |
| RX 9070, Vulkan | 3 164 t/s | ≈ 1,3 s |
The difference between these cards is therefore measured in fractions of a second on a short document. It becomes noticeable when the prompt reaches tens of thousands of tokens, as in an entire code repository or a batch of PDFs, or when each agent request rereads large contexts. For low-volume use, choose the card based on its cost per GB of VRAM, not this criterion.
#Compared with RX 9070: same VRAM, RDNA 4
The RX 9070 is the other current-generation 16 GB card. It has 640 GB/s of bandwidth versus 576 GB/s, a 220 W TDP versus 260 W, and 3,584 stream processors. Under Vulkan, it generates at 119.71 tokens per second and reads at 3,164: 35% faster at prompt reading, and nearly tied for generation. The practical difference lies elsewhere: ROCm, Ollama, and Windows behave differently for these two generations, as detailed on the RX 9070 page. One concrete example: Ollama’s documentation lists the RX 9070 among the ROCm cards under Linux but not under Windows, where it uses Vulkan, while the RX 7900 GRE appears in both systems’ lists. If you work under Windows and require ROCm, this favors the GRE; if you accept Vulkan, the difference disappears.
| Criterion | RX 7900 GRE | RX 9070 |
|---|---|---|
| Bandwidth | 576 GB/s | 640 GB/s |
| Card power | 260 W | 220 W |
| pp512 processing | 2 336,31 t/s | 3 164,10 t/s |
| Generation tg128 | 116,11 t/s | 119,71 t/s |
#Getting started
- 01Choose the systemAMD supports Radeon RX 7000 only on Ubuntu 24.04.4, Ubuntu 22.04.5, RHEL 10.1, and RHEL 9.7. On Windows, Ollama uses a ROCm v7 or HIP7 stack.
- 02Install OllamaThe ROCm v7 driver is required on Linux. Vulkan, enabled by default, serves as a fallback and performs well on this card.
- 03Load a model that fitsFirst test a 12B or 20B, then a 24B, with a short context. Check with ollama ps that the model is entirely on the GPU.
- 04Compare the two backendsRun llama-bench with Vulkan and then with ROCm, with and without Flash Attention. On this card, Flash Attention slowed generation under Vulkan in publicly reported measurements.
#What to use the 7900 GRE for
What sets it apart is the compute-to-memory ratio: plenty of compute power for 16 GB and 576 GB/s. It fits cases where time to first token matters as much as write speed, such as an assistant that ingests long documents or a RAG pipeline whose prompt contains thousands of tokens. For a short chat, the advantage disappears, and the 7800 XT, with more bandwidth, performs just as well or better.
| Usage | Preferred card | Reason |
|---|---|---|
| Chat, text generation | RX 7800 XT or RX 9070 | Higher bandwidth, equally fast or faster generation |
| RAG, long documents | RX 9070, then RX 7900 GRE | Faster prompt processing |
| 30 to 32B models | None of the three | You need 20 to 24 GB of VRAM |
| Tight used-market budget | The cheapest of the three | The cost of a GB of VRAM on /prix-gpu-ia makes the difference |
In summary, for the three 16 GB cards: the 7800 XT is the straightforward choice for conversation, the 9070 is the fastest for prompt processing and the most power-efficient at 220 W, and the 7900 GRE is the opportunistic choice when its price drops below the other two. None unlocks 30-billion-parameter models, which require more memory.
#2026 verdict
- A sensible choice
- If you find the GRE priced at or below a 7800 XT, and your prompts are long.
- A less relevant choice
- If it costs more than a 7800 XT and you mainly do conversation: you're paying for compute that generation doesn't use.
- Check before buying
- Make sure the system correctly detects 16 GB of VRAM and the card's correct model, using rocminfo on Linux or Task Manager on Windows.
- Local AI graphics card prices, recorded twice a week
- Ollama documentation: supported AMD cards
- ROCm GPU compatibility, AMD documentation
- Vulkan scoreboard for the llama.cpp repository
#Frequently asked questions
Is the RX 7900 GRE good for local LLMs?+
RX 7900 GRE or RX 7800 XT for local AI?+
Which models can you run on a 16 GB RX 7900 GRE?+
Does the RX 7900 GRE work with Ollama?+
Vulkan or ROCm on RX 7900 GRE?+
When does the RX 7900 GRE become too limiting?+
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.