Which LLM can run on an AMD Radeon 6700XT?

The search “6700xt llm” comes up often among people with an RDNA 2-generation card who want a local model without buying another GPU. Good news for the 6700xt llm combination: this card’s 12 GB of GDDR6 is enough for 7- to 14-billion-parameter models at 4-bit quantization, with comfortable throughput for chat, coding, and document summarization. The main constraint is not compute power but video memory, and the second obstacle is software: RX 6700 XT is not officially supported by ROCm, which requires a tweak or the Vulkan backend. This article covers the relevant specs, VRAM requirements by quantization, realistic throughput, and licenses, then answers frequently asked questions before directing you to the configurator.

The RX 6700 XT versus local inference

On paper, the Radeon RX 6700 XT is a 2021 gaming card. For an LLM, three figures really matter.

For comparison, a 16 GB card like the one described in the Radeon LLM guide RX 6800 XT opens up 24B models in Q4, and a 24 GB card like the RX 7900 XTX reaches 32B. The 6700 XT remains in the 12 GB class, the same as the RTX 3060 12 GB on the NVIDIA side. The guide dedicated to RX 6700 XT and the selection best LLM for Radeon RX 6700 XT list the catalog models that fit within this envelope.

VRAM consumed by quantization: Q4, Q5, Q8, FP16

The calculation is simple: number of parameters × bytes per parameter, plus the KV cache, which grows with context length. The orders of magnitude below assume a context of 8,192 tokens and an 8-bit KV cache; they are estimated from common GGUF files.

The tipping point is clear: above 12 GB, llama.cpp and Ollama offload the remaining layers to the processor, and throughput drops by a factor of 5 to 10. The guide: “which LLM for 12 GB of VRAM” details the size × quantization × context combinations that remain 100% on the GPU.

What the top of the catalog won't run

Part of the quelllm.fr catalog is beyond the reach of a 6700 XT, and it's better to say so clearly than to let people hope for a software miracle.

The interesting edge case is MoE models of around 120B with few active parameters. Qwen 3.5 122B-A10B weighs ~73 GB at Q4 under the Apache 2.0 license, Qwen3.8 Flash Next 125B-A6B ~72 GB with only 6B active, and Mistral Small 4 ~72 GB under Apache 2.0. With 64 GB of system RAM and the card's 12 GB, llama.cpp can load these models by offloading the experts to the CPU. The resulting throughput is only a few tokens per second, estimated at 3 to 8 depending on the RAM and CPU, to be confirmed on your configuration. It's usable for batch processing overnight, not for interactive chat. The weights are published on the Alibaba page on Hugging Face et the Mistral page on Hugging Face.

Actual tokens/sec and software choice

On RDNA 2, throughput depends as much on the backend as on the model. Two approaches work.

Orders of magnitude measured by the community on a 6700 XT, to be confirmed on your machine:

For a chat interface on top of Ollama, Open WebUI deploys in a container and manages history, documents, and multiple models.

Licenses and Real-World Use Cases

The license determines what you can do with the outputs and whether the model can be used in a product. The catalog displays the license for each entry; the major families available at the 12 GB size are Apache 2.0, MIT, Llama Community, and a few proprietary licenses. The license filter on the catalog lets you immediately rule out licenses with commercial-use restrictions.

What a 6700 XT does well every day:

To compare scores across models, the Open LLM Leaderboard remains the public benchmark, with the usual caveat: a benchmark measures a specific task, not your use case.

Should I upgrade my card?

If you run into the 12 GB limit, three tiers are available on the AMD side.

Le comparator places two cards side by side to show what the additional VRAM actually unlocks for a given model.

FAQ

Q: Is RX 6700 XT compatible with ROCm?

Not officially. AMD does not list the gfx1031 among the supported ROCm targets. In practice, Ollama and llama.cpp compiled for ROCm work when forcing HSA_OVERRIDE_GFX_VERSION=10.3.0, which presents the card as a gfx1030. If this setting causes problems, llama.cpp's Vulkan backend offers comparable throughput without depending on ROCm.

Q: What is the maximum model size on 12 GB?

A 14B in Q4 or Q5 remains entirely in VRAM with 8k context. A 24B or 27B requires an aggressive Q3 or partial CPU offloading, with a sharp drop in throughput. 30B MoE models with 3B active parameters are the best compromise beyond 14B, provided the experts are offloaded to the processor.

Q: Can you run DeepSeek R1 671B or Qwen 3 235B on a 6700 XT?

No. DeepSeek R1 671B requires ~400 GB in Q4, Qwen 3 235B-A22B ~142 GB. No combination of RAM + 12 GB VRAM on a desktop PC is enough. MoE models of around 120B with 6 to 10B active parameters are the absolute limit, with 64 GB of RAM and a throughput of a few tokens per second.

Q: Windows or Linux for an LLM on a 6700 XT?

Both work. On Windows, llama.cpp's Vulkan backend and recent Ollama builds eliminate the need to install ROCm. On Linux, ROCm with the override variable provides slightly faster prompt processing. The difference in generation throughput remains around 10%, estimated, and does not by itself justify changing operating systems.

Q: How much additional system RAM should you plan for?

32 GB is enough for models that fit in VRAM and for a 30B-A3B MoE with experts on the CPU. Moving to 64 GB is only worthwhile if you want to attempt approximately 120B MoEs with full offloading, such as Qwen3.8 Flash Next 125B-A6B, at the cost of lower throughput.

Q: Is a long context possible on 12 GB?

Yes, with trade-offs. An 8B model’s KV cache uses about 1 GB per 8k-token increment at 16-bit; quantized to 8-bit, an 8B model in Q4 reaches 32k context within 12 GB. A 14B model is limited to 16k under the same conditions. The catalog pages indicate the maximum context supported by each model.

Conclusion

For the 6700xt llm query, the answer fits in one sentence: a 7- to 14-billion-parameter model in Q4 or Q5, loaded through Ollama with the ROCm override or through llama.cpp in Vulkan, runs entirely in VRAM at a comfortable throughput. MoE models with offloaded experts extend the range, while the top of the catalog remains out of reach. To get the exact list of models compatible with your 12 GB, your RAM, and your target license, use the quelllm.fr configurator.

The hardware for running an LLM locally

To run these models comfortably locally, a RTX 5070 Ti offers an excellent price/performance ratio. Compare prices:

Amazon GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395) →

Affiliate links — BestLLMfor may earn a commission from purchases at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.

Article published and updated on by Mohamed Meguedmi · Data source: /api/models.json · Content license: CC BY 4.0.

An error or update to report? Contribute.