Which LLM on Radeon RX 7900 XTX (24 GB) ?
The RX 7900 XTX (24 GB GDDR6, up to 960 GB/s according to AMD) is an excellent card for local LLMs under Ollama or llama.cpp: it handles any model whose weights remain below approximately 20 GB in Q4, such as a dense 27B or a 30–35B MoE. It is supported by ROCm 7; beyond 24 GB, the model spills into RAM.
The 7900 XTX dates back to 2022, but its 24 GB still make it a reference card on the AMD side. This guide explains what it can run, with which software, at what speed according to public benchmarks, and when an RTX or a newer card is preferable.
Choosing a machine? Our picks by budget →
Buying alternative for this guide: Radeon RX 9070 XT 16GB (ASUS Prime OC).
Why this choice? Our complete guide on Radeon RX 9070 XT 16GB (ASUS Prime OC) →
Compare all options by budget, from €800 to €3,500 →
On the go: which laptop for local AI →
Affiliate links — possible commission at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.
#RX 7900 XTX for LLMs: what 24 GB changes
The Radeon RX 7900 XTX runs a local LLM without difficulty as soon as the model fits within its 24 GB of VRAM: a 7B to 9B model in Q4 is very comfortable, a 27B model in Q4 (about 16 GB of weights) leaves room for context, and a 30- to 35-billion-parameter MoE model with 3 billion active parameters still fits, just barely. Two conditions: use Ollama, LM Studio, or llama.cpp, all compatible with its drivers, and stay under 24 GB, because beyond that the model spills into RAM and speed collapses. It has the largest VRAM capacity in AMD's consumer RDNA 3 lineup.
The manufacturer provides the figures that matter for local AI: 24 GB of GDDR6 and memory bandwidth of up to 960 GB/s. Bandwidth is the most useful figure here because text generation reads most of the model weights for each token: the higher it is, the higher the ceiling for tokens per second.
| Feature | Value |
|---|---|
| Memory | 24 GB GDDR6 |
| Bandwidth | up to 960 GB/s |
| Stream processors / compute units | 6 144 / 96 |
| AI accelerators | 192 |
| Typical card power | 355 W |
| Recommended minimum power supply | 800 W |
| Listed matrix formats | FP16 (123 TFLOPs), INT8, INT4; no FP8 |
#Drivers and software: what's compatible in 2026
For an AMD card, the first question is the software stack. The 7900 XTX appears in AMD's Radeon documentation ROCm compatibility matrix, version 7.2.1, for Ubuntu 22.04, 24.04, and 25.10. Ollama, for its part, states that it requires the AMD ROCm v7 driver on Linux and lists the 7900 XTX for both Linux and Windows. Two practical consequences: on Linux, installation goes through the amdgpu-install utility; on Windows, the card is recognized without any special steps.
If ROCm causes problems, there is a second path. Ollama specifies that support for additional GPUs under Windows and Linux goes through Vulkan, enabled by default when the backend is installed. Vulkan is also llama.cpp's fallback backend: the llama.cpp guide for Vulkan explains the build process. For vLLM, the current documentation lists Radeon RX 7900 (gfx1100 and gfx1101) with ROCm 6.3 or later, and offers prebuilt packages for ROCm 7.0 and 7.2.1 with Python 3.12.
#Install Ollama on a 7900 XTX (Linux)
- 01Install the ROCm driverFollow AMD's ROCm documentation and use the amdgpu-install utility, as Ollama requires, then add your user to the render and video groups and restart.
- 02Verify that the card is detectedRun rocminfo and then rocm-smi: the 7900 XTX should appear with 24 GB. If nothing appears, Ollama will fall back to the processor.
- 03Install OllamaUse the official script or the Ollama installation guide on Linux, then download a model that fits in 24 GB.
- 04Control placementAfter an initial response, ollama ps shows the portion of the model placed on the GPU. It should display 100% GPU; otherwise, the model or context is too large.
#Which models fit in 24 GB, and how fast
The complete list of models by memory tier is in the “Which LLM for 24 GB of VRAM” guide; here is the key information for this card. The weights below come from the QuelLLM catalog, in Q4, excluding context. The theoretical ceiling is bandwidth divided by the weights read for each token: it cannot be exceeded and is not reached in practice.
| Model | Q4 weight | Margin on 24 GB | Theoretical ceiling |
|---|---|---|---|
| Qwen 3.5 9B | 6 GB | 18 GB | 160 tokens/s |
| gpt-oss 20B (MoE) | 13 GB | 11 GB | depends on the active weights |
| Devstral Small 2 24B | 14 GB | 10 GB | ≈ 68 tokens/s |
| Qwen 3.8 27B | 16 GB | 8 GB | 60 tokens/s |
| Granite 4.1 30B | 17 GB | 7 GB | ≈ 56 tokens/s |
| Qwen3-Coder 30B-A3B (MoE) | 19 GB | 5 GB | depends on the active weights |
| Qwen 3.6 35B-A3B (MoE) | 21 GB | 3 GB | depends on the active weights |
There are two ways to read this table. First, headroom: on 24 GB, a 21 GB model leaves only 3 GB for the KV cache and system, meaning a short context. The Qwen 3.8 27B, with 8 GB of headroom, is more comfortable for a long context. Second, MoEs: a 35B-A3B model reads only its 3 billion active parameters for each token, so it generates much faster than a dense 27-billion-parameter model, but it still has to fit entirely in VRAM.
#What public benchmarks measure on this card
The llama.cpp community reference table measures, on ROCm and with the same model (Llama 2 7B in Q4_0), prompt-processing speed (pp512) and generation speed (tg128). On the 7900 XTX, one participant recorded 3552 tokens/s for prompt processing and 167 tokens/s for generation without Flash Attention, then 3874 and 170 with Flash Attention. The table's author warns that results vary by driver, system, and card manufacturer, even with the same chip.
| Card | AMD bandwidth | Prompt pp512 | Generation tg128 |
|---|---|---|---|
| RX 7900 XTX | 960 GB/s | 3 552 t/s | 167 t/s |
| RX 9070 XT | 640 GB/s | 5 055 t/s | 101 t/s |
| RX 9060 XT | 320 GB/s | 1 420 t/s | 68 t/s |
This table shows two things. Generation follows memory bandwidth: the 7900 XTX (960 GB/s) generates about 1.65 times faster than the 9070 XT (640 GB/s) and 2.5 times faster than the 9060 XT (320 GB/s). Prompt processing, on the other hand, depends more on matrix compute power, and the newer-generation 9070 XT outperforms the 7900 XTX in this respect. For chat with long responses, bandwidth and 24 GB matter most; for RAG with large input documents, prompt-processing speed matters more.
#70B, 30B MoE, and dual cards: how far to go
A dense 70B model in Q4 weighs about 40 GB in weights according to the site's benchmark, nearly twice the VRAM: more than 16 GB would have to be placed in RAM, and speed would then be limited by RAM and the PCIe bus, not the card. We do not provide a throughput figure for this case because we lack a verifiable source; just consider it impractical. MoE models with 30 to 35 billion parameters are, at the same general quality, the practical answer to this 24 GB limit.
Two 7900 XTX cards change the picture for capacity. In llama.cpp, the default split mode divides the model by layers across the cards, with pipeline operation, and an option lets you choose the proportions for each card. You then have 48 GB, enough for a 70B in Q4 with some context. The gain is in what fits, not in tokens-per-second speed: each token passes through both cards one after the other. On the hardware side, two 355 W cards represent 710 W before the rest of the PC, which requires a properly sized power supply, two suitable PCIe slots, and a ventilated case.
#Context and KV cache: the real limit of 24 GB
A model's weight is only part of the memory it uses: the KV cache grows with context length. Ollama uses a default context window of 4,096 tokens according to its documentation, configurable with the OLLAMA_CONTEXT_LENGTH variable; recent models advertise 128,000 to 262,000 tokens, but accessing them requires additional memory. There are two levers: enable Flash Attention, which Ollama uses automatically when the backend and card allow it, and quantize the KV cache with OLLAMA_KV_CACHE_TYPE (f16 by default, q8_0 to reduce memory). The guide to KV-cache quantization covers these settings in detail.
#7900 XTX, RTX 3090, or 4090: how to decide
On paper, the RTX 3090 also offers 24 GB, in GDDR6X according to NVIDIA. The choice therefore comes down to the software stack, environment, and current price, which we do not lock in here: check the graphics card price tracker for local AI before buying.
| Your situation | Preferred card | Why |
|---|---|---|
| Linux, Ollama or llama.cpp, tight budget | RX 7900 XTX | 24 GB supported by ROCm 7 and Vulkan |
| Windows, tools that require CUDA (some fine-tuning projects) | RTX 3090 or 4090 | CUDA remains the default target for most projects |
| You want vLLM in production | Check your model | vLLM lists the 7900 cards with ROCm; support depends on the version |
| Long context and large MoE | All 24 GB | VRAM matters more than the brand |
| Requires more than 24 GB | RTX 5090 (32 GB) or dual card | No 24 GB card can run a 70B model in Q4 |
If you are choosing between a new and used card, remember that the 7900 XTX dates from December 2022: the manufacturer’s warranty has often expired on a used card, and a card that has been used for mining or run continuously deserves a temperature check before purchase.
- Which LLM for 24 GB of VRAM: the complete model list
- Ollama with an AMD GPU (ROCm): detailed configuration
- LLM on RX 7900 XT (20 GB)
- LLM on RTX 3090 (24 GB)
- Quantize the KV cache to save VRAM
- Local AI graphics card prices
- Source: AMD’s spec sheet for the Radeon RX 7900 XTX
- Source: Ollama GPU documentation
- Source: llama.cpp performance table on ROCm
- Source: installing vLLM on a GPU
#Frequently asked questions
Is the RX 7900 XTX good for local LLMs?+
Which model should you choose on RX 7900 XTX?+
Can you use two RX 7900 XTX together?+
Does the RX 7900 XTX support FP8?+
ROCm or Vulkan on RX 7900 XTX?+
Does the 7900 XTX work with vLLM?+
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.