Intermediate 11 minRadeon RX 7000

Which LLM on Radeon RX 7900 XTX (24 GB) ?

Direct response

The RX 7900 XTX (24 GB GDDR6, up to 960 GB/s according to AMD) is an excellent card for local LLMs under Ollama or llama.cpp: it handles any model whose weights remain below approximately 20 GB in Q4, such as a dense 27B or a 30–35B MoE. It is supported by ROCm 7; beyond 24 GB, the model spills into RAM.

The 7900 XTX dates back to 2022, but its 24 GB still make it a reference card on the AMD side. This guide explains what it can run, with which software, at what speed according to public benchmarks, and when an RTX or a newer card is preferable.

Choosing a machine? Our picks by budget →

By Mohamed Meguedmi·Update 2026-09-29·Tested on Windows, macOS, and Linux
Recommended hardware

Buying alternative for this guide: Radeon RX 9070 XT 16GB (ASUS Prime OC).

Why this choice? Our complete guide on Radeon RX 9070 XT 16GB (ASUS Prime OC) →

Compare all options by budget, from €800 to €3,500 →

On the go: which laptop for local AI →

Affiliate links — possible commission at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.

#RX 7900 XTX for LLMs: what 24 GB changes

The Radeon RX 7900 XTX runs a local LLM without difficulty as soon as the model fits within its 24 GB of VRAM: a 7B to 9B model in Q4 is very comfortable, a 27B model in Q4 (about 16 GB of weights) leaves room for context, and a 30- to 35-billion-parameter MoE model with 3 billion active parameters still fits, just barely. Two conditions: use Ollama, LM Studio, or llama.cpp, all compatible with its drivers, and stay under 24 GB, because beyond that the model spills into RAM and speed collapses. It has the largest VRAM capacity in AMD's consumer RDNA 3 lineup.

The manufacturer provides the figures that matter for local AI: 24 GB of GDDR6 and memory bandwidth of up to 960 GB/s. Bandwidth is the most useful figure here because text generation reads most of the model weights for each token: the higher it is, the higher the ceiling for tokens per second.

RX 7900 XTX technical specifications (source: AMD)
FeatureValue
Memory24 GB GDDR6
Bandwidthup to 960 GB/s
Stream processors / compute units6 144 / 96
AI accelerators192
Typical card power355 W
Recommended minimum power supply800 W
Listed matrix formatsFP16 (123 TFLOPs), INT8, INT4; no FP8
i
Power supply: 800 W, not 850 W
AMD lists 800 W as the minimum recommended power supply for this card. Some partner card manufacturers recommend more: follow the specifications for the exact model you buy, and leave headroom if the processor is power-hungry.

#Drivers and software: what's compatible in 2026

For an AMD card, the first question is the software stack. The 7900 XTX appears in AMD's Radeon documentation ROCm compatibility matrix, version 7.2.1, for Ubuntu 22.04, 24.04, and 25.10. Ollama, for its part, states that it requires the AMD ROCm v7 driver on Linux and lists the 7900 XTX for both Linux and Windows. Two practical consequences: on Linux, installation goes through the amdgpu-install utility; on Windows, the card is recognized without any special steps.

If ROCm causes problems, there is a second path. Ollama specifies that support for additional GPUs under Windows and Linux goes through Vulkan, enabled by default when the backend is installed. Vulkan is also llama.cpp's fallback backend: the llama.cpp guide for Vulkan explains the build process. For vLLM, the current documentation lists Radeon RX 7900 (gfx1100 and gfx1101) with ROCm 6.3 or later, and offers prebuilt packages for ROCm 7.0 and 7.2.1 with Python 3.12.

#Install Ollama on a 7900 XTX (Linux)

  1. 01
    Install the ROCm driver
    Follow AMD's ROCm documentation and use the amdgpu-install utility, as Ollama requires, then add your user to the render and video groups and restart.
  2. 02
    Verify that the card is detected
    Run rocminfo and then rocm-smi: the 7900 XTX should appear with 24 GB. If nothing appears, Ollama will fall back to the processor.
  3. 03
    Install Ollama
    Use the official script or the Ollama installation guide on Linux, then download a model that fits in 24 GB.
  4. 04
    Control placement
    After an initial response, ollama ps shows the portion of the model placed on the GPU. It should display 100% GPU; otherwise, the model or context is too large.
Checks
rocminfo | head -30
rocm-smi
ollama run gpt-oss:20b
ollama ps

#Which models fit in 24 GB, and how fast

The complete list of models by memory tier is in the “Which LLM for 24 GB of VRAM” guide; here is the key information for this card. The weights below come from the QuelLLM catalog, in Q4, excluding context. The theoretical ceiling is bandwidth divided by the weights read for each token: it cannot be exceeded and is not reached in practice.

Q4 weights (QuelLLM catalog) and theoretical ceiling at 960 GB/s
ModelQ4 weightMargin on 24 GBTheoretical ceiling
Qwen 3.5 9B6 GB18 GB160 tokens/s
gpt-oss 20B (MoE)13 GB11 GBdepends on the active weights
Devstral Small 2 24B14 GB10 GB≈ 68 tokens/s
Qwen 3.8 27B16 GB8 GB60 tokens/s
Granite 4.1 30B17 GB7 GB≈ 56 tokens/s
Qwen3-Coder 30B-A3B (MoE)19 GB5 GBdepends on the active weights
Qwen 3.6 35B-A3B (MoE)21 GB3 GBdepends on the active weights

There are two ways to read this table. First, headroom: on 24 GB, a 21 GB model leaves only 3 GB for the KV cache and system, meaning a short context. The Qwen 3.8 27B, with 8 GB of headroom, is more comfortable for a long context. Second, MoEs: a 35B-A3B model reads only its 3 billion active parameters for each token, so it generates much faster than a dense 27-billion-parameter model, but it still has to fit entirely in VRAM.

!
No made-up tokens/s
The throughput figures you find online depend on the model, quantization, context, and software version. The only comparable public reference we cite is the llama.cpp community table below, using a small model. For your models, measure with llama-bench or with the statistics returned by Ollama.

#What public benchmarks measure on this card

The llama.cpp community reference table measures, on ROCm and with the same model (Llama 2 7B in Q4_0), prompt-processing speed (pp512) and generation speed (tg128). On the 7900 XTX, one participant recorded 3552 tokens/s for prompt processing and 167 tokens/s for generation without Flash Attention, then 3874 and 170 with Flash Attention. The table's author warns that results vary by driver, system, and card manufacturer, even with the same chip.

Llama 2 7B Q4_0, llama.cpp ROCm (community table, without Flash Attention)
CardAMD bandwidthPrompt pp512Generation tg128
RX 7900 XTX960 GB/s3 552 t/s167 t/s
RX 9070 XT640 GB/s5 055 t/s101 t/s
RX 9060 XT320 GB/s1 420 t/s68 t/s

This table shows two things. Generation follows memory bandwidth: the 7900 XTX (960 GB/s) generates about 1.65 times faster than the 9070 XT (640 GB/s) and 2.5 times faster than the 9060 XT (320 GB/s). Prompt processing, on the other hand, depends more on matrix compute power, and the newer-generation 9070 XT outperforms the 7900 XTX in this respect. For chat with long responses, bandwidth and 24 GB matter most; for RAG with large input documents, prompt-processing speed matters more.

i
Small model, big gap from reality
A 7B in Q4_0 weighs far less than the 24 GB models in the previous section. Do not infer the speed of a 27B from these figures: at the same bandwidth, it will be several times slower.

#70B, 30B MoE, and dual cards: how far to go

A dense 70B model in Q4 weighs about 40 GB in weights according to the site's benchmark, nearly twice the VRAM: more than 16 GB would have to be placed in RAM, and speed would then be limited by RAM and the PCIe bus, not the card. We do not provide a throughput figure for this case because we lack a verifiable source; just consider it impractical. MoE models with 30 to 35 billion parameters are, at the same general quality, the practical answer to this 24 GB limit.

Two 7900 XTX cards change the picture for capacity. In llama.cpp, the default split mode divides the model by layers across the cards, with pipeline operation, and an option lets you choose the proportions for each card. You then have 48 GB, enough for a 70B in Q4 with some context. The gain is in what fits, not in tokens-per-second speed: each token passes through both cards one after the other. On the hardware side, two 355 W cards represent 710 W before the rest of the PC, which requires a properly sized power supply, two suitable PCIe slots, and a ventilated case.

#Context and KV cache: the real limit of 24 GB

A model's weight is only part of the memory it uses: the KV cache grows with context length. Ollama uses a default context window of 4,096 tokens according to its documentation, configurable with the OLLAMA_CONTEXT_LENGTH variable; recent models advertise 128,000 to 262,000 tokens, but accessing them requires additional memory. There are two levers: enable Flash Attention, which Ollama uses automatically when the backend and card allow it, and quantize the KV cache with OLLAMA_KV_CACHE_TYPE (f16 by default, q8_0 to reduce memory). The guide to KV-cache quantization covers these settings in detail.

Reduced context and KV cache
OLLAMA_CONTEXT_LENGTH=32768 OLLAMA_KV_CACHE_TYPE=q8_0 ollama serve

#7900 XTX, RTX 3090, or 4090: how to decide

On paper, the RTX 3090 also offers 24 GB, in GDDR6X according to NVIDIA. The choice therefore comes down to the software stack, environment, and current price, which we do not lock in here: check the graphics card price tracker for local AI before buying.

Decision criteria for 24 GB
Your situationPreferred cardWhy
Linux, Ollama or llama.cpp, tight budgetRX 7900 XTX24 GB supported by ROCm 7 and Vulkan
Windows, tools that require CUDA (some fine-tuning projects)RTX 3090 or 4090CUDA remains the default target for most projects
You want vLLM in productionCheck your modelvLLM lists the 7900 cards with ROCm; support depends on the version
Long context and large MoEAll 24 GBVRAM matters more than the brand
Requires more than 24 GBRTX 5090 (32 GB) or dual cardNo 24 GB card can run a 70B model in Q4

If you are choosing between a new and used card, remember that the 7900 XTX dates from December 2022: the manufacturer’s warranty has often expired on a used card, and a card that has been used for mining or run continuously deserves a temperature check before purchase.

→
The 20 GB and 16 GB alternatives
If 24 GB is beyond your budget, the RX 7900 XT (20 GB) and the RX 9070 XT (16 GB) each have their own guide. The 9070 XT is newer, with FP8 listed by AMD, but has 8 GB less.

#Frequently asked questions

Frequently asked questions
Is the RX 7900 XTX good for local LLMs?+
Yes, mainly thanks to its 24 GB of VRAM and 960 GB/s of bandwidth according to AMD. It runs models of up to about 27 billion parameters in Q4 without spilling over, as well as MoE models with 30 to 35 billion parameters. It is compatible with Ollama, llama.cpp, and LM Studio, under ROCm 7 or Vulkan.
Which model should you choose on RX 7900 XTX?+
For general use, a dense 27-billion-parameter model such as Qwen 3.8 27B in Q4, with about 16 GB of weights, leaves 8 GB for context. For speed, use an MoE such as gpt-oss 20B or Qwen 3.6 35B-A3B. For coding, use Devstral Small 2 24B or Qwen3-Coder 30B-A3B. Always check with ollama ps that everything is on the GPU.
Can you use two RX 7900 XTX together?+
Yes, with llama.cpp, whose default mode distributes the layers and KV cache across the cards. You get 48 GB, enough for a 70B model in Q4, but no additional tokens-per-second performance because the cards work in a pipeline. Plan for more than 710 W for the two cards, meaning a power supply well above 1,000 W.
Does the RX 7900 XTX support FP8?+
Not according to AMD's specifications, which list FP16, INT8, and INT4 for matrix computation on this card, but no FP8. Radeon RX 9000, based on the RDNA 4 architecture, list FP8. For most uses with Q4 or Q5 quantized GGUF models, this has no direct impact.
ROCm or Vulkan on RX 7900 XTX?+
Start with ROCm, Ollama's official path on Linux with the ROCm v7 driver. If installation gets stuck, Vulkan, enabled by default in Ollama, lets you get started without ROCm, generally with performance that varies by driver. Compare both with llama-bench on your own model before deciding.
Does the 7900 XTX work with vLLM?+
Yes, according to the vLLM documentation, which lists Radeon RX 7900 (gfx1100 and gfx1101) with ROCm 6.3 or later and offers prebuilt packages for ROCm 7.0 and 7.2.1. However, it is more demanding to install than Ollama, and is mainly useful if you serve multiple users at once.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.