Best LLM for Ryzen AI 9
Choose the best LLM for Ryzen AI 9 is not a matter of raw power, but of matching the available unified memory, the model architecture (dense or Mixture-of-Experts), and the actual throughput of the RDNA iGPU and XDNA NPU. On these platforms, a MoE model with a low number of active parameters often outperforms a smaller dense model. This article details VRAM requirements by quantization, expected throughput, licenses, benchmark reference points, and concrete use cases for building a consistent local inference workstation on Ryzen AI 9.
Understanding the Ryzen AI 9 memory constraint
Ryzen AI 9 APUs (HX and Max) use unified memory shared between the CPU, iGPU, and NPU. The amount that can be allocated to the GPU directly determines which model you can load. On a Ryzen AI Max 395 configuration with 128 GB of unified memory, the usable inference capacity far exceeds that of an HX 370 limited to approximately 32 GB.
Sizing guidelines (orders of magnitude, to be confirmed based on your BIOS allocation):
- Q4 : ~0.55 to 0.60 GB of VRAM per billion stored parameters
- Q5 : approximately +20% compared with Q4
- Q8 : about twice the Q4
- FP16 : about four times that of Q4
For the complete NPU/iGPU allocation process, see the Ryzen AI 9 HX guide and the Ryzen AI Max 395 guide. The quantization choice is detailed in our Q4/Q5/Q8 guide.
Why favor low-activation MoE models
On an iGPU with moderate memory bandwidth, the limiting factor is the number of parameters actifs per token, not the total. A Mixture-of-Experts model activates only a fraction of its weights on each pass.
- Qwen3.8 Flash Next 125B-A6B (125B, ~72 GB Q4, ctx 256000): only ~6B active parameters, making it an efficient candidate for interactive throughput on large unified memory. See the Qwen3.8 Flash Next profile and the editor Alibaba on Hugging Face.
- Qwen 3.5 122B-A10B (122B, ~73 GB Q4, ctx 262000, Apache 2.0): ~10B active, a good quality/throughput tradeoff. See the spec sheet Qwen 3.5 122B-A10B.
- Mixtral 8x22B Instruct (141B, ~82 GB Q4, ctx 64000, Apache 2.0): a historic MoE from Mistral AI on Hugging Face, a safe bet. See the Mixtral 8x22B model card.
These three models only fit in Q4 on a high-unified-memory configuration (such as Ryzen AI Max 395, 128 GB). On an HX 370 with 32 GB, they are out of reach locally without very slow disk offloading.
Compact dense models and alternatives
If your workload requires a more predictable dense model:
- Mistral Small 4 (119B, ~72 GB Q4, ctx 256000, Apache 2.0): permissive license for commercial use. See the Mistral Small 4 entry.
- Nemotron 3 Super 120B (120B, ~72 GB Q4, ctx 128000, NVIDIA Open Model License): see the Nemotron 3 Super page et NVIDIA on Hugging Face.
- DBRX Instruct (132B, ~76 GB Q4, ctx 32768, Databricks Open Model License): see the DBRX Instruct sheet.
- dots.llm1 Instruct (142B, ~85 GB Q4, ctx 32768, MIT): see the dots.llm1 spec sheet.
With the same VRAM, a dense model uses more bandwidth per token than a low-activation MoE: expect lower throughput on an iGPU. Precise tokens/sec throughput on Ryzen AI depends on the backend and memory allocation, and still needs to be confirmed on your machine.
Licensing and compliance
For professional use, the license matters just as much as throughput:
- Apache 2.0 : Qwen 3.5 122B-A10B, Mixtral 8x22B, Mistral Small 4 — broad commercial use.
- MIT : dots.llm1 Instruct — highly permissive.
- NVIDIA Open Model License : Nemotron 3 Super 120B — check the attribution clauses.
- Databricks Open Model License : DBRX Instruct — specific conditions to review.
For European regulatory considerations, refer to our AI Act compliance guide.
Recommended inference backends
The choice of engine strongly affects the throughput achieved:
- llama.cpp : Vulkan/ROCm support for AMD iGPUs, quantized GGUF format, ideal for self-hosting on Ryzen AI.
- Ollama : a practical layer for managing GGUF models locally.
- vLLM : server-throughput oriented, relevant if you expose an internal API.
Stay strictly local for inference: the goal of a Ryzen AI 9 workstation is to keep your data on the machine.
FAQ
Q: Which model fits on a Ryzen AI 9 HX 370 (32 GB)?
With ~32 GB of unified memory, some of which is reserved for the GPU, the models listed here (68 GB+ in Q4) will not fit. The HX 370 is better suited to lighter models outside this selection. See the catalog filtered by VRAM and the configurator for sizing suited to your actual memory.
Q: MoE or dense for an AMD iGPU?
A low-activation MoE such as Qwen3.8 Flash Next 125B-A6B (~6B active) uses less bandwidth per token than a 120B dense model. On an iGPU with limited bandwidth, the MoE generally delivers better interactive throughput at comparable quality, at the cost of a larger total memory footprint.
Q: Which quantization should you choose?
Q4 maximizes the number of loadable parameters and remains the usual balance point. Q5 slightly improves fidelity for ~20% more VRAM. Q8 doubles the footprint: reserve it for 128 GB configurations. Details in the Q4/Q5/Q8 guide.
Q: Does the NPU accelerate LLM inference?
The XDNA NPU is primarily designed for specific workloads, and its integration with open-source LLM backends is evolving. In practice, the RDNA iGPU via Vulkan/ROCm handles most inference today. Check the current support in llama.cpp — the exact status remains to be confirmed depending on your driver.
Q: Which license is required for commercial use?
Favor Apache 2.0 (Qwen 3.5 122B-A10B, Mixtral 8x22B, Mistral Small 4) or MIT (dots.llm1 Instruct) for the greatest freedom of use. The NVIDIA and Databricks licenses impose specific conditions that you must read before any deployment.
Conclusion
Le best LLM for Ryzen AI 9 depends above all on your unified memory: on a 128 GB configuration such as the Ryzen AI Max 395, low-activation MoE models like Qwen 3.5 122B-A10B or Qwen3.8 Flash Next 125B-A6B offer the best throughput/quality ratio in Q4, while Mistral Small 4 remains a permissive dense option. Check your actual throughput before committing to a choice. Size your system with the configurator or browse the catalog complet.
The hardware for running an LLM locally
To run these models comfortably locally, a RTX 5070 Ti offers an excellent price/performance ratio. Compare prices:
Affiliate links — BestLLMfor may earn a commission from purchases at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.