Which LLM on Radeon RX 7700 XT (12 GB) ?
The Radeon RX 7700 XT (12 GB GDDR6, 432 GB/s) is officially supported by ROCm and by Ollama, on both Linux and Windows. It runs 8- to 14-billion-parameter models in Q4 without compromise, but stops short of the 20- to 24B models, which require 13 to 14 GB. No throughput is published for it on the llama.cpp scoreboards: its performance falls between that of a RX 6700 XT and a RX 7800 XT.
The RX 7700 XT is valuable mainly because of what its 12 GB enables and where it sits in the lineup: faster and better supported than the previous generation, but stuck below the 16 GB threshold that opens the door to 20B to 24B models. Here is what you need to decide whether it is enough, with cited measurements and clearly labeled estimates.
Choosing a machine? Our picks by budget →
Buying alternative for this guide: Radeon RX 9070 XT 16GB (ASUS Prime OC).
Why this choice? Our complete guide on Radeon RX 9070 XT 16GB (ASUS Prime OC) →
Compare all options by budget, from €800 to €3,500 →
On the go: which laptop for local AI →
Affiliate links — possible commission at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.
#What a RX 7700 XT does in 2026
A RX 7700 XT is suitable for a mid-sized local LLM: its 12 GB can accommodate Llama 3.1 8B, Gemma 4 12B, or Phi-4 14B in Q4 with room for context. Its 432 GB/s bandwidth, 12.5% higher than a RX 6700 XT, sets the generation ceiling: about 113 tokens per second on the scoreboards' reference 7B, versus 163 for a RX 7800 XT. It has a clear advantage over the 6700 XT: ROCm and Ollama officially list it, including under Windows. Its drawback is capacity, which rules it out for 20- to 24-billion-parameter models, where the 7800 XT's 16 GB becomes valuable.
- Architecture
- RDNA 3, a Navi 32 chip assembled from chiplets (1 GCD and 3 MCDs), launched on September 6, 2023.
- Compute
- 3,456 stream processors, 54 compute units.
- Memory
- 12 GB GDDR6, 192-bit bus, 432 GB/s.
- Power consumption
- 245 W of board power.
- Pricing
- Not listed here: check /prix-gpu-ia for the lowest price recorded each week.
#ROCm, Ollama, and Windows: official support
This is what distinguishes it from the previous generation. AMD’s compatibility table lists the RX 7700 XT under RDNA 3, target gfx1101, like the RX 7800 XT. The Ollama documentation lists it among the supported cards on Linux and Windows. One nuance affects the choice of operating system: according to AMD, the Radeon RX 7000 and 9000 are supported only on Ubuntu 24.04.4, Ubuntu 22.04.5, RHEL 10.1, and RHEL 9.7. Another distribution may work, but without guarantees. Vulkan remains the safety net: Ollama enables it by default when the backend is installed.
| Environment | Status | Source |
|---|---|---|
| ROCm on Linux (Ubuntu 24.04.4, 22.04.5, RHEL 10.1, 9.7) | Official, targets gfx1101 | AMD compatibility chart |
| Ollama on Linux | List of supported cards | Ollama documentation, ROCm v7 required |
| Ollama on Windows | List of supported cards | Ollama documentation, ROCm v7 / HIP7 stack |
| Vulkan (Ollama, llama.cpp) | General-purpose alternative | Enabled by default in Ollama |
#What fits in 12 GB of VRAM
On a 12 GB card, allow about 11 GB for the model and its KV cache if it isn't driving the display. The listed weights come from the QuelLLM catalog; the KV cache is added on top and grows with conversation length.
| Model | Weights in Q4 | Headroom for context | Verdict |
|---|---|---|---|
| Llama 3.1 8B | ≈ 6 GB | ≈ 5 GB | Very comfortable |
| Gemma 4 12B | ≈ 7 GB | ≈ 4 GB | Comfortable, long context possible |
| Phi-4 14B | ≈ 9 GB | ≈ 2 GB | Fits, context must be limited |
| gpt-oss 20B | ≈ 13 GB | none | Overflows, with partial offload to the processor |
| Devstral Small 2 24B | ≈ 14 GB | none | You need at least 16 GB |
To go beyond 14B without changing cards, there are two options: quantize the KV cache to recover one or two gigabytes, or choose a mixture-of-experts (MoE) model in which only a few parameters are active per token. The occupied memory remains the model’s full size, so a 20B MoE does not fit any better in 12 GB, but it generates faster if it fits.
- Which LLM for 12 GB of VRAM, across all GPUs
- Quantize the KV cache to reclaim VRAM
- VRAM calculator for your model and context
#Throughput: what we know and what we estimate
Let’s be precise about the evidence: the two public discussions in the llama.cpp repository, one for Vulkan and the other for ROCm, which list AMD cards using the same model (Llama 2 7B in Q4_0), contain no entry for the RX 7700 XT at the time of writing. We did not measure anything ourselves. What we can do is place it between two measured cards: the RX 6700 XT at 83.88 tokens per second in generation (384 GB/s bandwidth) and the RX 7800 XT at 118.27 (624 GB/s), both using Vulkan.
Bandwidth-based reasoning gives us an upper bound: 432 GB/s divided by 3.82 GB of weights yields about 113 tokens per second on the reference 7B model. The AMD cards in the same discussion reach between 59% and 83% of that (83% for the 6700 XT, 72% for the 7800 XT). On that basis, a value between 65 and 95 tokens per second seems plausible for the 7700 XT: this is a working hypothesis, not a result, to be confirmed with llama-bench.
| Model (Q4) | Weights | Ceiling at 432 GB/s | Estimated at 60–80% of the ceiling |
|---|---|---|---|
| Llama 3.1 8B | ≈ 6 GB | ≈ 72 tokens/s | ≈ 43 to 58 tokens/s |
| Gemma 4 12B | ≈ 7 GB | ≈ 62 tokens/s | ≈ 37 to 49 tokens/s |
| Phi-4 14B | ≈ 9 GB | ≈ 48 tokens/s | ≈ 29 to 38 tokens/s |
For conversation, anything above roughly ten tokens per second reads comfortably, so these estimates remain well above the practical threshold, even with a significant margin of error.
#Compared with its direct neighbors: 6700 XT, 7800 XT, and RTX 4070
The choice often comes down to three closely matched cards. Measurements published under Vulkan help position the RTX 4070, one of the 12 GB Ada-generation cards discussed here: 92.29 tokens per second for generation and 3,179 for prompt processing. Its 12 GB, like the 7700 XT’s, limits both cards to the same models; the difference comes down to the software ecosystem, prompt processing, power efficiency, and price. The RX 7700 XT’s throughput remains to be measured; the other three have published figures.
| Card | VRAM | Bandwidth | pp512 processing | Generation tg128 |
|---|---|---|---|---|
| RX 6700 XT | 12 GB | 384 GB/s | 1 051,20 t/s | 83,88 t/s |
| RX 7700 XT | 12 GB | 432 GB/s | unpublished | unpublished |
| RTX 4070 | 12 GB | uncited | 3 179,37 t/s | 92,29 t/s |
| RX 7800 XT | 16 GB | 624 GB/s | 2 017,33 t/s | 118,27 t/s |
How to read this table: moving from the 7700 XT to the 7800 XT doesn’t just add speed; it adds 4 GB of VRAM, giving you access to 20B to 24B models. This is the only criterion that truly changes what you can run. If you use an 8B or 12B model, the 7700 XT isn’t a limitation; if you’re targeting 24B coding models, it is.
- Radeon RX 7800 XT, the same chip with 16 GB
- Radeon RX 6700 XT, the previous generation with 12 GB
- RTX 4070: the NVIDIA alternative at 12 GB
#Getting started on Linux
- 01Check the systemMake sure your distribution is one of Ubuntu 24.04.4, Ubuntu 22.04.5, RHEL 10.1, or RHEL 9.7, the only versions AMD lists for this card.
- 02Install the ROCm driverOllama requires the ROCm v7 driver on Linux. Install it with the amdgpu-install utility described in AMD's documentation.
- 03Install Ollama, then a modelDownload Gemma 4 12B in Q4_K_M or an 8B model, then run it with ollama run.
- 04Check that the GPU is workingRun ollama ps: the processor column should indicate execution on the GPU. If the card is not detected, add the ollama user to the render group.
- 05Measure and compareRun llama-bench with the GGUF from the public discussion to compare your card with others' measurements and publish yours.
On Windows, the Ollama documentation specifies a ROCm v7 or HIP7 stack for Radeon RX 7000; installation is limited to updating the Radeon driver and then running Ollama. The guide on Ollama and AMD GPUs explains the environment variables useful for selecting a card when the system has several.
#When 12 GB is no longer enough
The threshold is clear: above 11 GB of weights, the model no longer fits entirely in VRAM. Ollama and llama.cpp then offload some layers to the processor, and speed collapses because system memory is much slower than GDDR6. Three situations come up repeatedly: a 24B coding model, a long context on a 14B model, or a second model loaded at the same time for RAG. Each one is solved with VRAM, not computing power.
| Your usage | Is 12 GB enough? | What changes next |
|---|---|---|
| Chat, summarization, and translation with an 8B to 12B model | Yes | None |
| Code assistant with a 14B, short context | Yes | Monitor the KV cache |
| 24B coding model or gpt-oss 20B | No | Take at least 16 GB |
| RAG with generation, embeddings, and reranker loaded together | Just | Choose 16 GB to avoid back-and-forth transfers |
| 32,000 tokens or more of context on a 14B | No | Quantize the KV cache or get 16 GB |
Before buying it used, check three specific points. The card’s 245 W power draw requires a suitable power supply and two properly connected PCIe power connectors. The exact model matters little for an LLM because all manufacturers use the same 12 GB of memory; only the cooling quality changes, which matters during long generation sessions. Finally, test the card with a real 9–10 GB model before approving the purchase: it’s the only test that uses most of the 12 GB of memory and the continuous weight rereading an LLM performs, which has nothing to do with a video game.
#2026 verdict
- If you already have it
- Keep it: it covers everything up to 14B with official support and no workaround.
- If you buy used
- Compare its price with that of a RX 7800 XT on /prix-gpu-ia: if the difference is modest, the 7800 XT’s additional 4 GB is worth it.
- If you are targeting 24B models
- Go straight to 16 GB or more: capacity, not throughput, is the limiting factor.
- Local AI graphics card prices, recorded twice a week
- Ollama documentation: supported AMD GPU
- AMD ROCm compatibility table
- Vulkan scoreboard for the llama.cpp repository
#Frequently asked questions
Is the RX 7700 XT good for local LLMs?+
Which models can you run on 12 GB with a RX 7700 XT?+
How many tokens per second does a RX 7700 XT generate?+
Does the RX 7700 XT work with Ollama on Windows?+
RX 7700 XT or RX 7800 XT for local AI?+
Do you need Ubuntu to use RX 7700 XT with ROCm?+
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.