Which LLM can run on an AMD Radeon 6700XT?
The search “6700xt llm” comes up often among people with an RDNA 2-generation card who want a local model without buying another GPU. Good news for the 6700xt llm combination: this card’s 12 GB of GDDR6 is enough for 7- to 14-billion-parameter models at 4-bit quantization, with comfortable throughput for chat, coding, and document summarization. The main constraint is not compute power but video memory, and the second obstacle is software: RX 6700 XT is not officially supported by ROCm, which requires a tweak or the Vulkan backend. This article covers the relevant specs, VRAM requirements by quantization, realistic throughput, and licenses, then answers frequently asked questions before directing you to the configurator.
The RX 6700 XT versus local inference
On paper, the Radeon RX 6700 XT is a 2021 gaming card. For an LLM, three figures really matter.
- VRAM : 12 GB of GDDR6, 192-bit bus.
- Memory bandwidth : 384 GB/s. This is the figure that caps throughput in tokens/sec, because each generated token rereads all active weights.
- Architecture : RDNA 2, identifier
gfx1031, 40 compute units, with no dedicated matrix cores.
For comparison, a 16 GB card like the one described in the Radeon LLM guide RX 6800 XT opens up 24B models in Q4, and a 24 GB card like the RX 7900 XTX reaches 32B. The 6700 XT remains in the 12 GB class, the same as the RTX 3060 12 GB on the NVIDIA side. The guide dedicated to RX 6700 XT and the selection best LLM for Radeon RX 6700 XT list the catalog models that fit within this envelope.
VRAM consumed by quantization: Q4, Q5, Q8, FP16
The calculation is simple: number of parameters × bytes per parameter, plus the KV cache, which grows with context length. The orders of magnitude below assume a context of 8,192 tokens and an 8-bit KV cache; they are estimated from common GGUF files.
- 7B to 8B model : Q4 ≈ 5 GB, Q5 ≈ 6 GB, Q8 ≈ 9 GB, FP16 ≈ 16 GB. Everything up to and including Q8 fits on 12 GB; FP16 does not.
- 12–14B model : Q4 ≈ 8 to 9 GB, Q5 ≈ 10 GB, Q8 ≈ 15 GB. Q4 and Q5 fit; Q8 spills into system RAM.
- 24B to 27B model : Q4 ≈ 15 to 17 GB. It does not fit entirely; Q3 at 12 GB is possible but degrades quality.
- 30B MoE model with 3B active : Q4 ≈ 18 GB in weight, but only the active experts are read for each token. With experts offloaded to the CPU, throughput remains usable.
The tipping point is clear: above 12 GB, llama.cpp and Ollama offload the remaining layers to the processor, and throughput drops by a factor of 5 to 10. The guide: “which LLM for 12 GB of VRAM” details the size × quantization × context combinations that remain 100% on the GPU.
What the top of the catalog won't run
Part of the quelllm.fr catalog is beyond the reach of a 6700 XT, and it's better to say so clearly than to let people hope for a software miracle.
- DeepSeek R1 671B : ~400 GB in Q4, MIT license. Impossible to run locally on a consumer PC, regardless of the card.
- Qwen 3 235B-A22B : ~142 GB in Q4, Apache 2.0. Out of reach even with 128 GB of RAM.
- DeepSeek V4 Flash 284B : ~170 GB in Q4, MIT. Reserved for multi-GPU servers or high-end Macs with unified memory.
- GLM 5.3 Flash 320B-A18B : ~186 GB in Q4, MIT. Same verdict.
The interesting edge case is MoE models of around 120B with few active parameters. Qwen 3.5 122B-A10B weighs ~73 GB at Q4 under the Apache 2.0 license, Qwen3.8 Flash Next 125B-A6B ~72 GB with only 6B active, and Mistral Small 4 ~72 GB under Apache 2.0. With 64 GB of system RAM and the card's 12 GB, llama.cpp can load these models by offloading the experts to the CPU. The resulting throughput is only a few tokens per second, estimated at 3 to 8 depending on the RAM and CPU, to be confirmed on your configuration. It's usable for batch processing overnight, not for interactive chat. The weights are published on the Alibaba page on Hugging Face et the Mistral page on Hugging Face.
Actual tokens/sec and software choice
On RDNA 2, throughput depends as much on the backend as on the model. Two approaches work.
- Ollama with ROCm : the 6700 XT is not on the official list, but the variable
HSA_OVERRIDE_GFX_VERSION=10.3.0makes it recognized as a card from the same family. The complete procedure is in the Ollama GPU AMD ROCm guide ; the repository Ollama on GitHub documents the environment variables. - llama.cpp with Vulkan : no specific driver required, works on Windows as well as Linux. The backend is maintained in the llama.cpp repository. Generation is close to ROCm, while prompt processing is slightly slower.
- vLLM : not a realistic option on RDNA 2; the vLLM documentation targets datacenter cards and recent RDNA 3 cards.
Orders of magnitude measured by the community on a 6700 XT, to be confirmed on your machine:
- 8B in Q4_K_M : 35 to 45 tokens/sec during generation.
- 14B in Q4_K_M : 18 to 25 tokens/sec.
- 30B-A3B MoE with experts on the CPU : 12 to 20 tokens/sec, estimated.
- 24B in partially offloaded Q4 : 4 to 8 tokens/sec, estimated.
For a chat interface on top of Ollama, Open WebUI deploys in a container and manages history, documents, and multiple models.
Licenses and Real-World Use Cases
The license determines what you can do with the outputs and whether the model can be used in a product. The catalog displays the license for each entry; the major families available at the 12 GB size are Apache 2.0, MIT, Llama Community, and a few proprietary licenses. The license filter on the catalog lets you immediately rule out licenses with commercial-use restrictions.
What a 6700 XT does well every day:
- Code assistant : a 7 to 14B model in Q4 responds quickly enough for completion and proofreading in an editor. The HumanEval and LiveCodeBench scores in the profiles help distinguish the candidates.
- Summarization and RAG on private documents : 8k to 16k context fits in VRAM with an 8B in Q5; beyond that, reduce KV-cache quantization.
- Chat in French : prioritize models whose documentation reports a multilingual MMLU score or a French evaluation, rather than the English-only MMLU score.
- Offline batch processing : classification, field extraction, reformulation, where throughput matters less than confidentiality.
To compare scores across models, the Open LLM Leaderboard remains the public benchmark, with the usual caveat: a benchmark measures a specific task, not your use case.
Should I upgrade my card?
If you run into the 12 GB limit, three tiers are available on the AMD side.
- 16 GB : used RX 6800 XT, RX 7800 XT, or new RX 9060 XT. The 9060 XT 16 GB benchmark shows what 4 GB more and RDNA 4 bring to 24B models.
- 20 GB : RX 7900 XT, which fits a 27B in Q5.
- 24 GB : RX 7900 XTX, the only consumer AMD card that loads a 32B model in Q4 with a usable context.
Le comparator places two cards side by side to show what the additional VRAM actually unlocks for a given model.
FAQ
Q: Is RX 6700 XT compatible with ROCm?
Not officially. AMD does not list the gfx1031 among the supported ROCm targets. In practice, Ollama and llama.cpp compiled for ROCm work when forcing HSA_OVERRIDE_GFX_VERSION=10.3.0, which presents the card as a gfx1030. If this setting causes problems, llama.cpp's Vulkan backend offers comparable throughput without depending on ROCm.
Q: What is the maximum model size on 12 GB?
A 14B in Q4 or Q5 remains entirely in VRAM with 8k context. A 24B or 27B requires an aggressive Q3 or partial CPU offloading, with a sharp drop in throughput. 30B MoE models with 3B active parameters are the best compromise beyond 14B, provided the experts are offloaded to the processor.
Q: Can you run DeepSeek R1 671B or Qwen 3 235B on a 6700 XT?
No. DeepSeek R1 671B requires ~400 GB in Q4, Qwen 3 235B-A22B ~142 GB. No combination of RAM + 12 GB VRAM on a desktop PC is enough. MoE models of around 120B with 6 to 10B active parameters are the absolute limit, with 64 GB of RAM and a throughput of a few tokens per second.
Q: Windows or Linux for an LLM on a 6700 XT?
Both work. On Windows, llama.cpp's Vulkan backend and recent Ollama builds eliminate the need to install ROCm. On Linux, ROCm with the override variable provides slightly faster prompt processing. The difference in generation throughput remains around 10%, estimated, and does not by itself justify changing operating systems.
Q: How much additional system RAM should you plan for?
32 GB is enough for models that fit in VRAM and for a 30B-A3B MoE with experts on the CPU. Moving to 64 GB is only worthwhile if you want to attempt approximately 120B MoEs with full offloading, such as Qwen3.8 Flash Next 125B-A6B, at the cost of lower throughput.
Q: Is a long context possible on 12 GB?
Yes, with trade-offs. An 8B model’s KV cache uses about 1 GB per 8k-token increment at 16-bit; quantized to 8-bit, an 8B model in Q4 reaches 32k context within 12 GB. A 14B model is limited to 16k under the same conditions. The catalog pages indicate the maximum context supported by each model.
Conclusion
For the 6700xt llm query, the answer fits in one sentence: a 7- to 14-billion-parameter model in Q4 or Q5, loaded through Ollama with the ROCm override or through llama.cpp in Vulkan, runs entirely in VRAM at a comfortable throughput. MoE models with offloaded experts extend the range, while the top of the catalog remains out of reach. To get the exact list of models compatible with your 12 GB, your RAM, and your target license, use the quelllm.fr configurator.
The hardware for running an LLM locally
To run these models comfortably locally, a RTX 5070 Ti offers an excellent price/performance ratio. Compare prices:
Affiliate links — BestLLMfor may earn a commission from purchases at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.