Best LLM for Ryzen AI 9

Choose the best LLM for Ryzen AI 9 is not a matter of raw power, but of matching the available unified memory, the model architecture (dense or Mixture-of-Experts), and the actual throughput of the RDNA iGPU and XDNA NPU. On these platforms, a MoE model with a low number of active parameters often outperforms a smaller dense model. This article details VRAM requirements by quantization, expected throughput, licenses, benchmark reference points, and concrete use cases for building a consistent local inference workstation on Ryzen AI 9.

Understanding the Ryzen AI 9 memory constraint

Ryzen AI 9 APUs (HX and Max) use unified memory shared between the CPU, iGPU, and NPU. The amount that can be allocated to the GPU directly determines which model you can load. On a Ryzen AI Max 395 configuration with 128 GB of unified memory, the usable inference capacity far exceeds that of an HX 370 limited to approximately 32 GB.

Sizing guidelines (orders of magnitude, to be confirmed based on your BIOS allocation):

For the complete NPU/iGPU allocation process, see the Ryzen AI 9 HX guide and the Ryzen AI Max 395 guide. The quantization choice is detailed in our Q4/Q5/Q8 guide.

Why favor low-activation MoE models

On an iGPU with moderate memory bandwidth, the limiting factor is the number of parameters actifs per token, not the total. A Mixture-of-Experts model activates only a fraction of its weights on each pass.

These three models only fit in Q4 on a high-unified-memory configuration (such as Ryzen AI Max 395, 128 GB). On an HX 370 with 32 GB, they are out of reach locally without very slow disk offloading.

Compact dense models and alternatives

If your workload requires a more predictable dense model:

With the same VRAM, a dense model uses more bandwidth per token than a low-activation MoE: expect lower throughput on an iGPU. Precise tokens/sec throughput on Ryzen AI depends on the backend and memory allocation, and still needs to be confirmed on your machine.

Licensing and compliance

For professional use, the license matters just as much as throughput:

For European regulatory considerations, refer to our AI Act compliance guide.

Recommended inference backends

The choice of engine strongly affects the throughput achieved:

Stay strictly local for inference: the goal of a Ryzen AI 9 workstation is to keep your data on the machine.

FAQ

Q: Which model fits on a Ryzen AI 9 HX 370 (32 GB)?

With ~32 GB of unified memory, some of which is reserved for the GPU, the models listed here (68 GB+ in Q4) will not fit. The HX 370 is better suited to lighter models outside this selection. See the catalog filtered by VRAM and the configurator for sizing suited to your actual memory.

Q: MoE or dense for an AMD iGPU?

A low-activation MoE such as Qwen3.8 Flash Next 125B-A6B (~6B active) uses less bandwidth per token than a 120B dense model. On an iGPU with limited bandwidth, the MoE generally delivers better interactive throughput at comparable quality, at the cost of a larger total memory footprint.

Q: Which quantization should you choose?

Q4 maximizes the number of loadable parameters and remains the usual balance point. Q5 slightly improves fidelity for ~20% more VRAM. Q8 doubles the footprint: reserve it for 128 GB configurations. Details in the Q4/Q5/Q8 guide.

Q: Does the NPU accelerate LLM inference?

The XDNA NPU is primarily designed for specific workloads, and its integration with open-source LLM backends is evolving. In practice, the RDNA iGPU via Vulkan/ROCm handles most inference today. Check the current support in llama.cpp — the exact status remains to be confirmed depending on your driver.

Q: Which license is required for commercial use?

Favor Apache 2.0 (Qwen 3.5 122B-A10B, Mixtral 8x22B, Mistral Small 4) or MIT (dots.llm1 Instruct) for the greatest freedom of use. The NVIDIA and Databricks licenses impose specific conditions that you must read before any deployment.

Conclusion

Le best LLM for Ryzen AI 9 depends above all on your unified memory: on a 128 GB configuration such as the Ryzen AI Max 395, low-activation MoE models like Qwen 3.5 122B-A10B or Qwen3.8 Flash Next 125B-A6B offer the best throughput/quality ratio in Q4, while Mistral Small 4 remains a permissive dense option. Check your actual throughput before committing to a choice. Size your system with the configurator or browse the catalog complet.

The hardware for running an LLM locally

To run these models comfortably locally, a RTX 5070 Ti offers an excellent price/performance ratio. Compare prices:

Amazon GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395) →

Affiliate links — BestLLMfor may earn a commission from purchases at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.

Article published and updated on by Mohamed Meguedmi · Data source: /api/models.json · Content license: CC BY 4.0.

An error or update to report? Contribute.