Intermediate 11 minRadeon RX 7000

Which LLM on Radeon RX 7900 GRE (16 GB) ?

Direct response

The Radeon RX 7900 GRE (16 GB GDDR6, 576 GB/s, 260 W) runs 20- to 24-billion-parameter models in Q4, like the RX 7800 XT, but without gaining generation speed: public measurements on Llama 2 7B report 116 tokens per second versus 118 for the 7800 XT under Vulkan. Its advantage lies elsewhere, in prompt processing under Vulkan, which is slightly faster (2,336 versus 2,017 tokens per second).

The 7900 GRE looks like a 7800 XT with stronger compute performance but less memory, and the benchmarks show it. This guide explains what its 16 GB can handle, where it differs from its two direct neighbors, the RX 7800 XT and RX 9070, and what use case makes it worth the price.

Choosing a machine? Our picks by budget →

By Mohamed Meguedmi·Update 2026-09-30·Tested on Windows, macOS, and Linux
Recommended hardware

Buying alternative for this guide: Radeon RX 9070 XT 16GB (ASUS Prime OC).

Why this choice? Our complete guide on Radeon RX 9070 XT 16GB (ASUS Prime OC) →

Compare all options by budget, from €800 to €3,500 →

On the go: which laptop for local AI →

Affiliate links — commission possible at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.

#What a RX 7900 GRE does in 2026

The RX 7900 GRE handles anything under 15 GB in Q4: Gemma 4 12B through Q8, gpt-oss 20B, and Devstral Small 2 24B with a short context. Its generation speed, capped by its 576 GB/s, is around 116 tokens per second on the benchmark scoreboards’ reference 7B model under Vulkan. That is the key takeaway: its 5,120 stream processors, versus 3,840 for the 7800 XT, speed up prompt processing but not generation, because it has neither more memory nor more bandwidth than its counterpart. AMD officially lists it in ROCm, and Ollama supports it on Linux and Windows.

Architecture
RDNA 3, Navi 31 (the RX 7900 XTX chip, reduced here), 80 compute units and 5,120 stream processors.
Memory
16 GB GDDR6 with a 256-bit bus, providing 576 GB/s.
Power consumption
260 W of board power.
Availability
Initially reserved for China (July 27, 2023), marketed internationally on February 27, 2024, at a launch price of $549.
Current price
See /prix-gpu-ia: the price changes every week.

#7900 GRE vs. 7800 XT: memory decides

Both cards offer 16 GB. The 7800 XT has 624 GB/s of bandwidth, while the 7900 GRE has 576 GB/s, or 8% less. During generation, when every token rereads all the weights, bandwidth is what matters: the 7800 XT therefore slightly outperforms the GRE in the llama.cpp Vulkan discussion, and more clearly with Flash Attention enabled. The GRE regains the advantage for prompt processing under Vulkan (16% more), where compute matters. Under ROCm, both public series show the 7800 XT ahead on both measurements.

RX 7900 GRE and RX 7800 XT (Llama 2 7B Q4_0, public llama.cpp measurements)
MetricRX 7900 GRERX 7800 XT
Vulkan, pp512 processing, without FA2 336,31 t/s2 017,33 t/s
Vulkan, tg128 generation, without FA116,11 t/s118,27 t/s
Vulkan, tg128 generation, with FA110,50 t/s124,86 t/s
ROCm, pp512 processing, without FA1 456,98 t/s2 151,81 t/s
ROCm, tg128 generation, without FA96,07 t/s100,94 t/s

These results come from different contributors, with different llama.cpp commits: the gap is an order of magnitude, not a laboratory comparison. They nevertheless agree that the GRE does not generate faster than a 7800 XT. If you find both at the same price, choose the 7800 XT for conversational use, and the GRE for use with long prompts.

#What fits in 16 GB of VRAM

Allow about 15 GB for the model and its KV cache when the card is not powering the display. The Q4 weights come from the QuelLLM catalog; the context is added on top.

What a 7900 GRE can handle (weights only)
ModelWeightsHeadroom for contextVerdict
Gemma 4 12B in Q4≈ 7 GB≈ 8 GBVery comfortable, long context
Gemma 4 12B in Q8≈ 13 GB≈ 2 GBFits, short context
gpt-oss 20B in Q4≈ 13 GB≈ 2 GBHolds up, short- to medium-length context
Devstral Small 2 24B in Q4≈ 14 GB≈ 1 GBToo cramped
Gemma 4 31B in Q4≈ 18 GBnoneDoesn't fit; you need 20 to 24 GB

The 16 GB boundary therefore lies between the 24B, which fits, and the 31B, which does not. To cross this boundary without moving to a 24 GB card, the RX 7900 XT at 20 GB provides an intermediate margin; the dedicated guide explains what it offers. For all 16 GB cards, regardless of brand, the VRAM-by-VRAM page compares the models.

#Measured throughput and bandwidth ceiling

None of the figures on this page are measurements of the site. The values come from public discussions in the llama.cpp repository, all of which use Llama 2 7B in Q4_0 (3.56 GiB) and llama-bench. Prompt processing (pp512) measures the ingestion speed for 512 tokens; generation (tg128) measures writing speed. For the RX 7900 GRE under Vulkan, the discussion reports 116.11 tokens per second during generation without Flash Attention and 110.50 with it. Under ROCm, it reports 96.07.

The theoretical generation ceiling, bandwidth divided by weight size, is 576 GB/s here for 3.82 GB, or about 151 tokens per second. The card reaches 77% under Vulkan without Flash Attention. Applying this efficiency to larger models gives a ballpark figure: it is not a measurement, and another driver or another version of llama.cpp will change it.

Estimate at 75% of the 576 GB/s ceiling (calculation, not measurement)
Model (Q4)WeightsTheoretical ceilingOrder of magnitude
Llama 3.1 8B≈ 6 GB≈ 96 tokens/s≈ 72 tokens/s
Gemma 4 12B≈ 7 GB≈ 82 tokens/s≈ 62 tokens/s
Phi-4 14B≈ 9 GB≈ 64 tokens/s≈ 48 tokens/s
Devstral Small 2 24B≈ 14 GB≈ 41 tokens/s≈ 31 tokens/s

#What reading the prompt changes in practice

Prompt processing seems abstract until you convert it into seconds of waiting. The calculation is simple: prompt token count divided by pp512 speed. It assumes that the speed measured on 512 tokens remains constant for a longer prompt, which is only approximately true, because processing slows as the context grows. The times below are therefore rough estimates for a document of about 4,000 tokens, or roughly ten pages.

Ingestion time for a 4,000-token prompt (calculated from pp512 measurements)
Card and backendpp512 speedTime to first token
RX 7900 GRE, Vulkan2 336 t/s≈ 1,7 s
RX 7800 XT, Vulkan2 017 t/s≈ 2,0 s
RX 7900 GRE, ROCm1 457 t/s≈ 2,7 s
RX 9070, Vulkan3 164 t/s≈ 1,3 s

The difference between these cards is therefore measured in fractions of a second on a short document. It becomes noticeable when the prompt reaches tens of thousands of tokens, as in an entire code repository or a batch of PDFs, or when each agent request rereads large contexts. For low-volume use, choose the card based on its cost per GB of VRAM, not this criterion.

#Compared with RX 9070: same VRAM, RDNA 4

The RX 9070 is the other current-generation 16 GB card. It has 640 GB/s of bandwidth versus 576 GB/s, a 220 W TDP versus 260 W, and 3,584 stream processors. Under Vulkan, it generates at 119.71 tokens per second and reads at 3,164: 35% faster at prompt reading, and nearly tied for generation. The practical difference lies elsewhere: ROCm, Ollama, and Windows behave differently for these two generations, as detailed on the RX 9070 page. One concrete example: Ollama’s documentation lists the RX 9070 among the ROCm cards under Linux but not under Windows, where it uses Vulkan, while the RX 7900 GRE appears in both systems’ lists. If you work under Windows and require ROCm, this favors the GRE; if you accept Vulkan, the difference disappears.

RX 7900 GRE and RX 9070 under Vulkan (Llama 2 7B Q4_0, without FA)
CriterionRX 7900 GRERX 9070
Bandwidth576 GB/s640 GB/s
Card power260 W220 W
pp512 processing2 336,31 t/s3 164,10 t/s
Generation tg128116,11 t/s119,71 t/s

#Getting started

  1. 01
    Choose the system
    AMD supports Radeon RX 7000 only on Ubuntu 24.04.4, Ubuntu 22.04.5, RHEL 10.1, and RHEL 9.7. On Windows, Ollama uses a ROCm v7 or HIP7 stack.
  2. 02
    Install Ollama
    The ROCm v7 driver is required on Linux. Vulkan, enabled by default, serves as a fallback and performs well on this card.
  3. 03
    Load a model that fits
    First test a 12B or 20B, then a 24B, with a short context. Check with ollama ps that the model is entirely on the GPU.
  4. 04
    Compare the two backends
    Run llama-bench with Vulkan and then with ROCm, with and without Flash Attention. On this card, Flash Attention slowed generation under Vulkan in publicly reported measurements.
Terminal
ollama ps
./llama-bench -m llama-2-7b.Q4_0.gguf -ngl 100 -fa 0,1

#What to use the 7900 GRE for

What sets it apart is the compute-to-memory ratio: plenty of compute power for 16 GB and 576 GB/s. It fits cases where time to first token matters as much as write speed, such as an assistant that ingests long documents or a RAG pipeline whose prompt contains thousands of tokens. For a short chat, the advantage disappears, and the 7800 XT, with more bandwidth, performs just as well or better.

Which 16 GB AMD card for which use case
UsagePreferred cardReason
Chat, text generationRX 7800 XT or RX 9070Higher bandwidth, equally fast or faster generation
RAG, long documentsRX 9070, then RX 7900 GREFaster prompt processing
30 to 32B modelsNone of the threeYou need 20 to 24 GB of VRAM
Tight used-market budgetThe cheapest of the threeThe cost of a GB of VRAM on /prix-gpu-ia makes the difference

In summary, for the three 16 GB cards: the 7800 XT is the straightforward choice for conversation, the 9070 is the fastest for prompt processing and the most power-efficient at 220 W, and the 7900 GRE is the opportunistic choice when its price drops below the other two. None unlocks 30-billion-parameter models, which require more memory.

#2026 verdict

A sensible choice
If you find the GRE priced at or below a 7800 XT, and your prompts are long.
A less relevant choice
If it costs more than a 7800 XT and you mainly do conversation: you're paying for compute that generation doesn't use.
Check before buying
Make sure the system correctly detects 16 GB of VRAM and the card's correct model, using rocminfo on Linux or Task Manager on Windows.

#Frequently asked questions

FAQ
Is the RX 7900 GRE good for local LLMs?+
Yes: its 16 GB loads 20- to 24-billion-parameter models in Q4, and it generates about 116 tokens per second on a 7B in Q4_0 under Vulkan, according to a public benchmark. Its advantage over the RX 7800 XT is limited to prompt processing.
RX 7900 GRE or RX 7800 XT for local AI?+
The 7800 XT for conversation: it has 624 GB/s of bandwidth versus 576 GB/s and generates slightly faster in public measurements. The 7900 GRE reads the prompt about 16% faster under Vulkan, which is useful for long documents. At the same price, the choice depends on this use case.
Which models can you run on a 16 GB RX 7900 GRE?+
Gemma 4 12B in Q4 or Q8, gpt-oss 20B in Q4, and Devstral Small 2 24B in Q4 fit, with a short context for the last two. A 31B model such as Gemma 4 in Q4, about 18 GB, exceeds the card's 16 GB.
Does the RX 7900 GRE work with Ollama?+
Yes, Ollama's documentation lists it under Linux with the ROCm v7 driver, and under Windows. AMD also lists it in ROCm, targeting gfx1100, but only under Ubuntu 24.04.4, 22.04.5, RHEL 10.1, and RHEL 9.7. Vulkan, enabled by default in Ollama, remains available as a fallback if detection fails.
Vulkan or ROCm on RX 7900 GRE?+
In public llama.cpp repository discussions, Vulkan generates faster than ROCm on this card (116 versus 96 tokens per second without Flash Attention) and also processes the prompt faster. Configurations vary from one series to another: test both with your model.
When does the RX 7900 GRE become too limiting?+
As soon as you target models with 30 billion parameters or more, or a very long context on a 24B, 16 GB is no longer sufficient. The RX 7900 XT with 20 GB or a 24 GB card are the next steps.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.

Prices in euros (€) are French market prices including VAT, as checked by BestLLMfor. US prices differ: the Amazon buttons show the current US price.