Beginner 11 minRadeon RX 6000

Which LLM on Radeon RX 6700 XT (12 GB) ?

Direct response

The RX 6700 XT (12 GB, 384 GB/s) runs models up to 14B in Q4 (about 9 GB) and a 12B model such as Gemma 4 12B comfortably, with room for context. It uses Vulkan rather than ROCm, so it appears in neither the official ROCm list nor Ollama’s list. During generation, it matches or exceeds a 12 GB RTX 3060; for prompt processing, it is nearly twice as slow.

This 2021 card was never designed for AI, but the software has not forgotten it either. This guide explains what it can run today, through which software stack, at what speed according to published measurements, and when replacement becomes more cost-effective than stubbornly persevering.

Choosing a machine? Our picks by budget →

By Mohamed Meguedmi·Update 2026-09-30·Tested on Windows, macOS, and Linux
Recommended hardware

Buying alternative for this guide: Radeon RX 9070 XT 16GB (ASUS Prime OC).

Why this choice? Our complete guide on Radeon RX 9070 XT 16GB (ASUS Prime OC) →

Compare all options by budget, from €800 to €3,500 →

On the go: which laptop for local AI →

Affiliate links — possible commission at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.

#What a RX 6700 XT does in 2026

A Radeon RX 6700 XT works very well as an entry-level card for a local LLM, provided you accept three facts: its 12 GB of memory can accommodate a 14-billion-parameter model in Q4 (about 9 GB of weights), but not a 24B; its reliable software path is Vulkan, through llama.cpp or Ollama, because AMD does not list it in ROCm; and its generation speed, limited by its 384 GB/s of bandwidth, is sufficient for a smooth conversation on an 8B to 12B model. On a 7B in Q4_0, a measurement published in the llama.cpp repository reports 83.88 tokens per second during generation under Vulkan, slightly faster than the RTX 3060 12 GB measured in the same discussion. The real weakness is prompt processing, which is twice as slow.

Architecture
RDNA 2, Navi 22 chip, launched on March 18, 2021 (Wikipedia page for the RX 6000 series).
Compute
2,560 stream processors, 40 compute units.
Memory
12 GB GDDR6, 192-bit bus, 384 GB/s.
Power consumption
230 W of board power.
Pricing
Not listed here: used prices change every week; check /prix-gpu-ia.

#Ollama, ROCm, or Vulkan: the path that works

The most common trap with this card is looking for a ROCm guide. The Ollama documentation lists Radeon RX cards from the 6800 to the 9070 XT for Linux: the 6700 XT is not included. On the ROCm side, AMD’s compatibility table lists only professional cards from RDNA 2, such as the PRO W6800 and V620. Ollama expands this scope with Vulkan, enabled by default when the backend is installed, and this is the path that remains open to a 6700 XT. A ROCm workaround exists—the HSA_OVERRIDE_GFX_VERSION variable forces a nearby LLVM target—but it is intended for experimentation.

Three software paths for a RX 6700 XT
PathStatus for this cardWhen to choose it
Vulkan (Ollama, llama.cpp, LM Studio)Works, publicly measured under llama.cppDefault choice on both Windows and Linux
Official ROCmCard absent from the ROCm list and from Ollama's listAvoid: no support guarantee
Forced ROCm (HSA_OVERRIDE_GFX_VERSION)Workaround documented by Ollama for other cardsOnly for testing, without expecting a proven gain
!
Don't count on ROCm for this card
Ollama states that ROCm does not support all AMD cards and suggests the HSA_OVERRIDE_GFX_VERSION variable as a workaround. For the 6700 XT, no source shows an improvement over Vulkan, whose results are published. Check your target with the rocminfo command before attempting anything.

#What fits in 12 GB of VRAM

The site's rule remains the same: the Q4 model weight plus the KV cache plus some headroom for the driver and display. On 12 GB, a realistic budget leaves about 11 GB for the model and its context if the card is not driving the display. The sizes below come from the QuelLLM catalog and our usual benchmarks; they are model weights only.

What a 12 GB card can handle (Q4, weights only)
ModelWeights in Q4Remaining contextVerdict
Llama 3.1 8B≈ 6 GB≈ 5 to 6 GBVery comfortable, long context possible
Gemma 4 12B≈ 7 GB≈ 4 to 5 GBComfortable, medium-to-long context
Phi-4 14B≈ 9 GB≈ 2 to 3 GBFits, context must be limited
gpt-oss 20B≈ 13 GBnegativeOverflows: partial fallback to the CPU
Devstral Small 2 24B≈ 14 GBnegativeToo large for 12 GB in Q4

The KV cache grows with the conversation length, which explains why a 14B model “fits” on paper but crawls as the history gets longer. The site's VRAM calculator gives the exact total for your model and context, and the guide to KV-cache quantization explains how to reclaim one or two gigabytes. For the full range of 12 GB cards, the dedicated page compares models without limiting itself to AMD.

#Measured throughput and bandwidth ceiling

None of the throughput figures on this page are measurements from the site. The figures come from the public llama.cpp repository discussion comparing cards under Vulkan with the same model, Llama 2 7B in Q4_0 (3.56 GiB), and the llama-bench command. Two values matter there: pp512, the prompt-reading speed for 512 tokens, and tg128, the generation speed for 128 tokens. For the RX 6700 XT, the discussion reports 1,051 tokens per second for reading and 83.88 for generation; with Flash Attention enabled, 1,011 and 81.86.

This result can be compared with a simple ceiling: during generation, each token requires the card to reread all the weights, so maximum speed is roughly bandwidth divided by model size. Here, 384 GB/s for 3.82 GB of weights gives a theoretical ceiling of 100 tokens per second; the card reaches 83% of it—a good result, better than that of several newer cards in the same discussion.

Estimate for common models (calculated, not measured)
Model (Q4)WeightsTheoretical ceiling: 384 GB/sRealistic order of magnitude (83%)
Llama 3.1 8B≈ 6 GB≈ 64 tokens/s≈ 50 to 55 tokens/s
Gemma 4 12B≈ 7 GB≈ 55 tokens/s≈ 45 tokens/s
Phi-4 14B≈ 9 GB≈ 43 tokens/s≈ 35 tokens/s

These ballpark figures assume that the performance measured on the 7B carries over to larger models, which is not guaranteed: a driver, a version of llama.cpp, or different quantization can shift the result by several points. For prompt processing, retain only the ratio between cards: that is what matters for RAG use with long documents.

#Against the RTX 3060 12 GB: generation wins, prompting loses

The usual comparison says that the RTX 3060 12 GB wins thanks to CUDA. Measurements published under Vulkan put that assumption into perspective. During generation, the RX 6700 XT outperforms the RTX 3060 in both modes; during prompt processing, the RTX 3060 performs almost twice as well. Both series come from the same discussion but different versions of llama.cpp: treat the gap as an order of magnitude, not as a definitive ranking.

RX 6700 XT and RTX 3060 under Vulkan (Llama 2 7B Q4_0, llama.cpp discussion)
MetricRX 6700 XTRTX 3060 (12 GB)
Prompt processing (pp512), without Flash Attention1 051,20 t/s1 815,70 t/s
Generation (tg128), without Flash Attention83,88 t/s75,94 t/s
Prompt processing (pp512), with Flash Attention1 010,90 t/s2 012,88 t/s
Generation (tg128), with Flash Attention81,86 t/s80,59 t/s

The practical consequence: for chatting with a model, the two cards are comparable; for querying a long document, the RTX 3060 responds faster to the first token. CUDA also retains its ecosystem advantage, with fewer workarounds to learn. At the same used-market price, the RTX 3060 12 GB remains the safer choice, and the 6700 XT is a good deal only if it’s already in your PC or significantly cheaper.

#Step-by-step setup

  1. 01
    Check the Vulkan driver
    On Windows, the Radeon driver includes Vulkan, and Ollama requires no additional steps. On Linux, install Mesa's Vulkan components (RADV) or the AMD driver, as documented by Ollama.
  2. 02
    Install Ollama or compile llama.cpp
    Ollama enables Vulkan by default once the backend is installed. With llama.cpp, compile with the -DGGML_VULKAN=on option.
  3. 03
    Give the card access under Linux
    Add the ollama user to the render group if the card isn't detected. For Ollama to read free VRAM, run it as root or grant it the capability specified in the command below.
  4. 04
    Run your first model
    Download an 8B or Gemma 4 12B in Q4_K_M, run it with the command ollama run, then check with ollama ps that the model is loaded on the GPU rather than the processor.
  5. 05
    Measure on your system
    Run llama-bench with the same GGUF as the public discussion to compare your card with its results before looking for a cause of slowdown.
Terminal
sudo setcap cap_perfmon+ep /usr/local/bin/ollama
# llama.cpp avec Vulkan
cmake .. -DGGML_VULKAN=on -DCMAKE_BUILD_TYPE=Release
./bin/llama-bench -m llama-2-7b.Q4_0.gguf -ngl 100 -fa 0,1

If throughput seems low under RADV, the llama.cpp discussion recommends launching with the RADV_PERFTEST=nogttspill environment variable, which fixes several performance issues. The other Vulkan settings are in the dedicated build guide.

#Limitations and when to move on

Three limitations stand out. First, memory: 12 GB rules out 20- to 32-billion-parameter models, where the quality of coding and reasoning assistants is determined. Next, prompt processing, which makes long-document sessions painful. Finally, support: a card outside the official list depends on Vulkan and llama.cpp updates being maintained, with no long-term guarantee.

When to stay, when to switch
Your needsStay on RX 6700 XT?Replacement option
Daily chat with an 8B to 12B modelYesNone
Local coding assistant with a 14BYes, short contextA 16 GB card if the context grows
20B to 30B models (gpt-oss 20B, Devstral 24B)No, insufficient memory16 to 20 GB card, dedicated page below
RAG on large documentsPossible, but slow to produce the first tokenCard with better prompt comprehension
Guaranteed official supportNoRadeon RX 7000 or 9000, RTX

#2026 verdict

It's worth it
If it is already in your machine: it is one of the few 2021 cards that can still run a 14B in Q4 with a generation yield above 80% of its theoretical ceiling.
Buy a RTX 3060 12 GB instead
If you’re starting from scratch and targeting the same budget: twice-as-fast prompt processing and zero software workarounds.
Switch cards
As soon as you want 20- to 30-billion-parameter models; prices change every week, and the /prix-gpu-ia page gives the cost per GB of VRAM.

#Frequently asked questions

FAQ
RX 6700 XT LLM: which models can you run?+
Models up to 14 billion parameters in Q4, or about 9 GB of weights, plus context: Llama 3.1 8B, Gemma 4 12B, Phi-4 14B. A 20B model such as gpt-oss (about 13 GB) or a 24B exceeds 12 GB and partially spills onto the CPU, causing a sharp drop in speed.
Does a 6700 XT work with Ollama?+
Yes, via Vulkan. Ollama does not list it in its ROCm list, which starts with RX 6800 on Linux, but it enables Vulkan by default when the backend is installed. Check with the ollama ps command that the model is loaded on the GPU.
Is ROCm viable on RX 6700 XT?+
Not officially: the card appears neither in AMD’s ROCm compatibility table, which lists only professional cards for RDNA 2, nor in the Ollama list. A workaround using the HSA_OVERRIDE_GFX_VERSION variable exists, but it remains experimental, and no public measurements show an advantage over Vulkan, whose results are published.
How many tokens per second on a RX 6700 XT?+
For Llama 2 7B in Q4_0, the public discussion in the llama.cpp repository reports 83.88 tokens per second for generation with Vulkan. For Gemma 4 12B in Q4, the bandwidth-to-weights ratio suggests about 45 tokens per second: a calculated estimate to verify on your system with llama-bench before citing it.
RX 6700 XT or RTX 3060 12 GB for an LLM?+
For generation, the two are comparable, with the Radeon even slightly ahead in measurements published under Vulkan. For prompt processing, the RTX 3060 is nearly twice as fast, and CUDA avoids workarounds. At the same used price, the RTX 3060 12 GB remains the simplest choice.
When should you give up on the RX 6700 XT for local AI?+
When you want 20- to 32-billion-parameter models, which 12 GB cannot hold, or when slow processing of long prompts interferes with your workflow. A 16- to 20-GB card, such as the Radeon RX 7800 XT or 7900 XT, is the logical next step.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.