Best LLM on Radeon RX 9070 XT (16 GB) in 2026
Identify the best RX 9070 XT LLM requires understanding two interconnected constraints: 16 GB of GDDR6 VRAM and AMD's ROCm ecosystem, now functional on the RDNA4 architecture. The RX 9070 XT represents in 2026 the most accessible entry point for running 27-billion-parameter models in Q4 quantization on consumer GPUs under Linux or Windows. This guide examines the hardware specs, selects the BestLLMfor catalog models suited to this card, compares licenses, provides speed benchmarks, and offers recommendations by use case.
RX 9070 XT and ROCm: what you need to know before choosing a model
The Radeon RX 9070 XT comes with 16 GB of GDDR6 memory on a 256-bit bus, for theoretical bandwidth of approximately 576 GB/s. The RDNA4 architecture has been supported since ROCm 6.x, which makes it possible to use llama.cpp compiled with the HIP backend, Ollama ≥ 0.3.x or LM Studio—provided you have recent amdgpu drivers on Linux, or a ROCm for Windows installation.
Key points before launching a model:
- Usable VRAM : ~15.5 GB in practice, after system and driver overhead
- Bandwidth : ~576 GB/s (estimated), versus ~1,008 GB/s for a RTX 4090 — inference speed is proportionally reduced
- Most stable backend : llama.cpp with
-DGGML_HIPBLAS=ONon Linux - Compatibility : check the AMD compatibility matrix for your exact kernel and driver
- Native accuracy : FP16 supported by RDNA4 compute units
The good news: models around 27B in Q4 quantization are perfectly suited to 16 GB, and the 24B family offers comfortable headroom for long contexts.
Which models fit in 16 GB of VRAM?
The rough rule for Q4 is about 0.55 GB per billion parameters. A 27B model therefore weighs ~15–16 GB in Q4, a 24B ~13–14 GB, and a 32B ~18–19 GB — excluding overhead.
26–27B models compatible with native Q4 (≤ 16 GB)
- Gemma 2 27B — Q4 VRAM : ~16 GB, ctx : 8 192
- Gemma 3 27B — Q4 VRAM : ~16 GB, ctx : 128 000
- Qwen 3.5 27B — Q4 VRAM : ~16 GB, ctx : 262 000
- Qwen 3.6 27B — Q4 VRAM : ~16 GB, ctx : 262 144
- Qwen 3.8 27B — Q4 VRAM : ~16 GB, ctx : 262 144
- Gemma 4 26B-A4B MoE — Q4 VRAM : ~16 GB, ctx : 128,000, mixture-of-experts architecture
- Trinity Mini 26B-A3B — Q4 VRAM : ~15 GB, ctx : 131 072
24B models (comfortable headroom, 2-3 GB free)
- Mistral Small 3.2 24B — Q4 VRAM : ~14 GB
- Magistral Small 24B — Q4 VRAM : ~14 GB
- Devstral Small 2 24B — Q4 VRAM : ~14 GB, ctx : 256 000
20–22B models (ample headroom, ~3–5 GB free)
- Codestral 22B v0.1 — Q4 VRAM : ~13 GB, ctx : 32 000
- EuroLLM 22B Instruct 2512 — Q4 VRAM : ~13 GB, European multilingual
- gpt-oss 20B — Q4 VRAM : ~13 GB, ctx : 128 000
Note on 32B models : in Q4 (~19 GB), they exceed the 16 GB available. In Q3 (~13-14 GB, or ~0.42 GB/B), it is technically possible with a measurable loss of accuracy. On the 9070 XT, sticking with 27B models in Q4 remains the most rational choice.
Top 5 recommended models for RX 9070 XT
1. Qwen 3.8 27B — versatility and long context
Qwen 3.8 27B is the most advanced version of the Qwen 3.x series in the 27B range, released by Alibaba under the Apache 2.0 license. Its 262,144-token context is one of the largest available in this size range. It fully fits in 16 GB at Q4 without overflow.
- License : Apache 2.0 (commercial use permitted)
- Q4 VRAM : ~16 GB
- Context : 262 144 tokens
- Strengths : general reasoning, coding, long-context questions
- Estimated speed : 20–30 tokens/s on RX 9070 XT (estimated from community benchmarks, proportional to bandwidth)
2. Gemma 3 27B — validated performance, 128k context
Gemma 3 27B from Google offers 128,000 tokens of context and a well-documented architecture. Its performance on the Hugging Face Open LLM Leaderboard place it among the leading models in the 27B category for understanding and following instructions.
- License : Gemma (restrictions for certain commercial uses — see Google’s Terms of Service)
- Q4 VRAM : ~16 GB
- Context : 128,000 tokens
- Strengths : long-range comprehension, precise instruction following
3. Magistral Small 24B — structured reasoning
Magistral Small 24B Mistral AI is designed for multi-step reasoning tasks and structured analysis. At ~14 GB in Q4, it leaves 2 GB of headroom on the 9070 XT — useful for maintaining an extended system context in production.
- License : Apache 2.0
- Q4 VRAM : ~14 GB
- Context : 128,000 tokens
- Strengths : analysis, structured summary, step-by-step agent, chain-of-thought reasoning
4. Devstral Small 2 24B — software development agent
Devstral Small 2 24B (Apache 2.0, Mistral AI) explicitly targets development agents with an exceptional 256,000-token context. It can ingest an entire codebase in a single pass, making it suitable for automated review, test generation, or documentation.
- License : Apache 2.0
- Q4 VRAM : ~14 GB
- Context : 256,000 tokens
- Strengths : code generation, refactoring, reading large codebases
5. Codestral 22B v0.1 — lightweight coding specialist
Codestral 22B v0.1 Mistral AI is one of the benchmarks in the under-23B coding category. It supports fill-in-the-middle (FIM) and achieves higher HumanEval scores than general-purpose models of the same size (exact scores on GGUF Q4: to be confirmed). It weighs only ~13 GB in Q4.
- License : Mistral Non-Production License (non-commercial use only)
- Q4 VRAM : ~13 GB
- Context : 32,000 tokens
- Strengths : code completion, FIM, IDE integration
Licenses: Apache 2.0, MIT, Gemma — what changes depending on usage
The license choice determines whether it can be used in a business or product.
Permissive licenses for commercial use (Apache 2.0 or MIT)
- Qwen 3.5 27B, Qwen 3.6 27B, Qwen 3.8 27B — Apache 2.0
- Trinity Mini 26B-A3B — Apache 2.0
- Magistral Small 24B, Devstral Small 2 24B, Mistral Small 3.2 24B — Apache 2.0
- EuroLLM 22B Instruct 2512, gpt-oss 20B — Apache 2.0
- Phi-4 14B — MIT
Restricted licenses
- License Gemma (Gemma 2 27B, Gemma 3 27B, Gemma 4 26B-A4B MoE): requires acceptance of the Google Gemma terms of service ; some commercial uses are subject to restrictions
- Mistral Non-Production License : Codestral 22B is restricted to noncommercial and research contexts
For a concise overview of the licenses available in the catalog, consult the open-weights LLM licensing guide.
Expected benchmarks and performance on RX 9070 XT
The scores below come from official publications or community benchmarks. Inference speeds on RX 9070 XT are estimated based on measurements from other GPUs weighted by their respective memory bandwidth.
MMLU — general evaluation
- Qwen 3.x 27B series: MMLU performance above 75% in reported 5-shot results on the page Qwen on HuggingFace for variants of this size (to be confirmed on Q4)
- Gemma 3 27B: results published by Google in the official Gemma documentation — verify degradation in Q4 depending on the quantization level
- Magistral Small 24B: reasoning metrics published by Mistral AI in the release notes
HumanEval — code generation
- Devstral Small 2 24B: designed for agent coding, with a HumanEval score higher than general-purpose models of the same size (to be confirmed on GGUF Q4)
- Codestral 22B: the sub-23B code category reference according to community benchmarks reported on r/LocalLLaMA
Inference speed (estimated)
- 27B Q4 models: 20-35 tokens/s on RX 9070 XT
- 24B Q4 models: 35–50 tokens/s
- 14–22B Q4 models: 50–80 tokens/s
These ranges depend on the backend (llama.cpp HIP vs native ROCm), the degree of parallelism, and the active context length. With a context near the maximum (262k tokens), speed drops significantly.
Concrete use cases
Software development and code review Devstral Small 2 24B for multi-file agents and large projects; Codestral 22B for completion and FIM in an IDE. Both run with headroom on the 9070 XT.
Long-document analysis Qwen 3.8 27B with its 262,144-token context can analyze lengthy contracts or reports in a single request. Gemma 3 27B (128,000 tokens) is a solid alternative under the Gemma license.
Multilingualism and French EuroLLM 22B Instruct 2512 is trained specifically on European languages, including French, German, and Spanish. It runs with approximately 3 GB of headroom on the 9070 XT.
Fast and lightweight model Phi-4 14B (MIT, Microsoft) uses ~9 GB in Q4 and leaves 7 GB for a very large context or simultaneous sessions. Relevant when speed matters more than model size.
Reasoning and agents Magistral Small 24B for structured chains of thought and simple agents. Trinity Mini 26B-A3B (Apache 2.0, Arcee AI) at ~15 GB in Q4 offers a good reasoning-to-size tradeoff.
To compare two models directly, use the BestLLMfor comparator.
FAQ
Q: Is RX 9070 XT well supported by ROCm?
ROCm 6.x includes support for the RDNA4 architecture (navi48). On Linux with recent amdgpu-dkms drivers, llama.cpp HIP works reliably. On Windows, ROCm for Windows is still maturing — see the AMD compatibility matrix before any configuration. Performance on Linux is generally better than in a Windows environment for now.
Q: Can you run a 32B model on a RX 9070 XT with 16 GB?
In Q4 (~19 GB for a 32B), no—the VRAM is exceeded. In Q3 (~13–14 GB), it's feasible with a noticeable loss of accuracy. In practice, 27B models in Q4 offer a better quality-to-size ratio on this card: the same inference level, full Q4 precision, and no risk of overflow.
Q: Which quantization should you choose: Q4, Q5, or Q8?
Q4_K_M is the best starting point: little quality loss and half the size of FP16. Q5_K_M improves accuracy for about +25% more VRAM. Q8_0 is close to FP16 but exceeds 16 GB for most 27B models. On the 9070 XT, Q4_K_M on 27B or Q5_K_M on 24B are rational choices. Full FP16 on 27B (~30 GB) is out of reach.
Q: Does Ollama work on the RX 9070 XT under Linux?
Yes. Ollama ≥ 0.3.x automatically detects compatible AMD GPUs under Linux if the amdgpu-dkms drivers are properly installed. No manual configuration is required in most cases. Under Windows, support depends on the ROCm for Windows version available at the time of installation.
Q: Which model do you recommend for everyday French?
EuroLLM 22B Instruct 2512 is the most targeted for European languages. Mistral Small 3.2 24B from Mistral AI (Apache 2.0) offers excellent performance in French for general tasks and assisted writing, with plenty of headroom on the 9070 XT.
Q: Qwen 3.8 27B or Gemma 3 27B: which should you choose?
Both use ~16 GB in Q4. Qwen 3.8 27B is under Apache 2.0 (free commercial use) and offers a larger context window (262,144 vs. 128,000 tokens). Gemma 3 27B is subject to the Gemma license, which is more restrictive for commercial use. If the license is not a factor, the respective scores on the Open LLM Leaderboard let you decide based on your specific task.
Conclusion
Le best RX 9070 XT LLM in 2026 falls in the 24-27B range: Qwen 3.8 27B and Gemma 3 27B for general versatility, Devstral Small 2 24B and Codestral 22B for software development, EuroLLM 22B for European multilingualism, Magistral Small 24B for structured reasoning. With 16 GB of RDNA4 VRAM and ROCm, this card delivers serious self-hosting performance. Explore the BestLLMfor complete catalog or use the configurator to refine your selection based on your GPU, target license, and use case.
The hardware for running an LLM locally
To run these models comfortably locally, a Radeon RX 9070 XT offers an excellent price/performance ratio. Compare prices:
Affiliate links — BestLLMfor may earn a commission on purchases, at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.