LLM performance on AMD Radeon 9070XT GPU
The query “9070xt llm” comes up frequently among owners of the Radeon RX 9070 XT who want to run a model locally without an NVIDIA card. Good news: with 16 GB of GDDR6 and ROCm support for the RDNA 4 architecture, a 9070xt llm setup is viable for most personal use cases, provided you understand its memory limitations. This page details the card specs relevant to inference, VRAM usage by quantization level (Q4, Q5, Q8, FP16), expected tokens/sec throughput, models from the BestLLMfor catalog that can run with CPU offloading, how to read the benchmarks, and the use cases and tools (llama.cpp, Ollama, vLLM) that actually work on AMD.
Technical specifications of the RX 9070 XT for inference
What matters for an LLM is not raw computing power but memory bandwidth and VRAM capacity. On these two points, the card ranks as follows:
- Architecture : RDNA 4, ROCm gfx1201 identifier
- VRAM : 16 GB GDDR6, 256-bit bus
- Bandwidth : approximately 640 GB/s (manufacturer specification)
- Compute units : 64 CUs, 4096 stream processors
- Typical consumption : 304 W under load
- Software support : ROCm 6.4 and later versions, with llama.cpp's Vulkan backend as an alternative
For context: bandwidth is close to that of a RX 7900 XT, but VRAM remains at 16 GB versus 20 or 24 GB on the previous generation. This is the trade-off to keep in mind before buying. The RX 9070 XT's page lists models ranked by compatibility, and the dedicated installation guide covers the setup step by step.
VRAM: what fits in 16 GB depending on quantization
Rule of thumb for estimating the memory used by a dense model: approximately 0.55 to 0.6 GB per billion parameters in Q4, 0.7 in Q5, 1.05 in Q8, and 2 in FP16, plus the KV cache, which grows with the context. With 16 GB, keeping 1 to 1.5 GB for the system and KV cache:
- Q4 (Q4_K_M) : dense models up to approximately 20 to 24 billion parameters, 8k to 16k context
- Q5 : up to approximately 16 to 18 billion parameters
- Q8 : up to approximately 12 to 13 billion parameters
- FP16 : up to approximately 7 billion parameters, rarely useful for local inference
The models in the reference catalog cited on this page are all much larger than that. Let's take three examples from the 118- to 128-billion-parameter range, where only the active experts work on each token:
- Qwen 3.5 122B-A10B : Q4 ~73 GB, Q5 ~88 GB (estimated), Q8 ~130 GB (estimated), FP16 ~245 GB (estimated). Spec sheet: Qwen 3.5 122B-A10B
- Mistral Small 4 (119B): Q4 ~72 GB, Q8 ~128 GB (estimated). Specs: Mistral Small 4
- Qwen3.8 Flash Next 125B-A6B : Q4 ~72 GB, only 6 billion active parameters. Profile: Qwen3.8 Flash Next
None of these models fits in 16 GB of VRAM. They are still interesting for a 9070 XT thanks to offloading experts to system RAM, as explained below.
Tokens/sec throughput: estimated order of magnitude
During generation, the throughput of a model that fits entirely in VRAM is bounded by memory bandwidth: each token requires reading all the weights. At 640 GB/s, the theoretical limit for an 8 GB model in Q4 is approximately 80 tokens/sec; in practice, 55% to 70% of that ceiling is observed. Rough figures, all estimated and to be confirmed with your own configuration:
- Dense model with 7 to 9 billion parameters, Q4, 100% in VRAM : 45 to 60 tokens/sec
- Dense model with 12 to 14 billion parameters, Q4 : 28 to 38 tokens/sec
- Dense model with 22 to 24 billion parameters, Q4, short context : 15 to 22 tokens/sec
- Dense model with 30 billion parameters or more, Q4 : partial offloading to the CPU, often 3 to 6 tokens/sec, tedious for daily use
For the catalog’s mixture-of-experts models, throughput depends mainly on system RAM. With 64 to 96 GB of dual-channel DDR5, attention layers in VRAM, and experts in RAM, you can target 8 to 15 tokens/sec (estimated) on a model with 10 billion active parameters such as Qwen 3.5 122B-A10B, and 12 to 20 tokens/sec (estimated) on Qwen3.8 Flash Next 125B-A6B. Prompt processing remains significantly slower than with everything in VRAM. The method is documented on the llama.cpp GitHub through tensor placement options on the CPU. The local benchmark comparator provides figures measured on other cards to calibrate your expectations.
Expert models that can be used with CPU offloading, and their licenses
For serious use on a 9070 XT with plenty of RAM, here are the candidates from the BestLLMfor catalog in the 118- to 142-billion-parameter range, with the license to be checked before any commercial use:
- Qwen 3.5 122B-A10B (Alibaba): Apache 2.0, 262k context, ~73 GB Q4 VRAM. Weights published on Alibaba's Hugging Face page
- Qwen3.8 Flash Next 125B-A6B (Qwen): “other, open weights” license; read carefully
- Mistral Small 4 (Mistral): Apache 2.0, 256k context. Weights on the Hugging Face page for Mistral
- Mistral Medium 3.5 128B : Modified MIT, Q4 VRAM ~74 GB. Details: Mistral Medium 3.5
- Nemotron 3 Super 120B (NVIDIA): NVIDIA Open Model License, Q4 VRAM ~72 GB. Weights on the Hugging Face page for NVIDIA
- Laguna S 2.1 (Poolside): OpenMDW 1.1, code-oriented, ~68 GB Q4 VRAM, the lightest of the bunch
- dots.llm1 Instruct (Rednote): MIT, ~85 GB Q4 VRAM, limited to 32k context
- Mixtral 8x22B Instruct (Mistral): Apache 2.0, ~82 GB Q4 VRAM, older architecture
Apache 2.0 and MIT allow commercial use without special clauses. The “Modified MIT,” “NVIDIA Open Model License,” and “OpenMDW” licenses add conditions that you should read in the model repository before building a product on it. If you are hesitating between several candidates, the comparator places two sheets side by side.
Benchmarks: how to read the scores for this GPU
Published scores (MMLU for general knowledge, HumanEval and LiveCodeBench for code, GPQA for scientific reasoning) measure the model at full precision, not your configuration. Three precautions are essential on a 9070 XT:
- Quantization slightly degrades scores : in Q4_K_M, the loss generally remains below 2 points on MMLU for models with more than 12 billion parameters, and is greater for smaller models. Figure to be confirmed model by model.
- The catalog scores should be cross-checked with the 'Open LLM Leaderboard, which evaluates open weights under consistent conditions.
- Truncated context changes the results : a model advertised with 256k context but run with 8k on 16 GB of VRAM won't behave as it does in long-context tests.
The right approach is to compare two models on your own tasks, using your prompts, rather than relying on an aggregate ranking.
Concrete use cases and compatible AMD tools
On 9070 XT, everyday use cases that work well are the coding assistant in the editor, document summarization and rewriting, translation, analysis of private files in local RAG, and tool-enabled agents for short tasks. Very long-context use cases, such as analyzing contracts several hundred pages long, require either a small model with a quantized KV cache or a second card.
On the software side:
- Ollama has supported ROCm on RDNA 4 since the 2025 versions; the procedure is in the Ollama guide on AMD GPUs, and the project is tracked on Ollama's GitHub
- llama.cpp offers two backends: ROCm (HIP) for the best throughput and Vulkan for easier installation, which is useful if ROCm fails to recognize the card
- vLLM targets multi-user service instead; its ROCm support is documented at the vLLM website but target professional cards first, and test them before relying on them
If loading fails, the GPU isn’t detected, or the system silently falls back to the CPU, the Ollama GPU troubleshooting guide lists the common causes. If you're considering adding a second card to exceed 16 GB, the multi-GPU llama.cpp guide explains the layer distribution.
FAQ
Q: Is the RX 9070 XT a good choice for a first local LLM setup?
Yes for models up to 24 billion parameters in Q4, with comfortable throughput and a lower price than NVIDIA cards with equivalent VRAM. The weak point remains the 16 GB: compare it with a used RX 7900 XTX with 24 GB if you are targeting larger models. The GPU selection guide details these trade-offs.
Q: What's the difference from the RX 9060 XT 16 GB for LLMs?
The same amount of VRAM, so the same models can be loaded, but the 9060 XT has a 128-bit bus and roughly half the bandwidth. Throughput in tokens/sec follows approximately the same ratio. The 9060 XT 16 GB benchmark provides concrete measurements to compare with this page's estimates.
Q: Can Qwen 3.5 122B-A10B run on a 9070 XT?
Not in VRAM alone: the Q4 version weighs about 73 GB. With 96 GB or more of system RAM, llama.cpp can keep the shared layers on the GPU and send the experts to RAM. Throughput then drops to around 8 to 15 tokens/sec (estimated), acceptable for conversational use but slow for long prompts.
Q: ROCm or Vulkan on this card?
ROCm generally delivers 10 to 25% more throughput (estimated) and better prompt processing. Vulkan installs without AMD-specific dependencies and works on nearly all Linux distributions and on Windows. Start with Vulkan to validate the hardware, then switch to ROCm if the improvement interests you and the installed version recognizes gfx1201.
Q: How much additional system RAM should you plan for?
To stay with models that fit in VRAM, 32 GB is enough. To run the catalog’s mixture-of-experts models with offloading, target at least 64 GB, ideally 96 to 128 GB of dual-channel DDR5. RAM speed directly affects tokens/sec in this mode, more than the processor itself.
Q: Which models run without a GPU if the card is occupied?
Models with a low number of active parameters remain usable on CPU alone, at reduced throughput. The page dedicated to GPU-free inference explains the thread and quantization settings that limit the damage, and the ranking CPU only lists the candidates.
Conclusion
A 9070xt llm setup makes sense if you want a fast local LLM for 7- to 24-billion-parameter models, and remains usable with the catalog’s mixture-of-experts models as long as you invest in system RAM. The throughput figures given here are estimates to be confirmed on your machine. To find the model suited to your 16 GB of VRAM, RAM, and workload, use the configurator or browse the full catalog filtered by VRAM.
The hardware for running an LLM locally
To run these models comfortably locally, a RTX 5070 Ti offers an excellent price/performance ratio. Compare prices:
Affiliate links — BestLLMfor may earn a commission on purchases, at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.