Best LLM for AMD AI PCs (Ryzen AI & NPU)

Un AMD AI PC refers to a machine equipped with a Ryzen AI processor—that is, an APU combining Zen 5 cores, an integrated Radeon GPU, and an XDNA 2 NPU. Choosing the best LLM for an AMD AI PC therefore comes down to a specific question: how much unified memory can the machine allocate to the model, and what bandwidth feeds that memory? The answer is not the same for a Ryzen AI 9 HX laptop with 32 GB and a Ryzen AI Max+ 395 workstation with 128 GB. This article covers both machine families, the models in the quelllm.fr catalog that actually fit in their memory, the role of the NPU, compatible inference tools, licenses, and answers the most frequently asked questions.

Two AMD AI PC families, two memory budgets

The term AMD AI PC covers very different machines. You need to distinguish between them before discussing models.

The practical rule is simple. The quantized model's size must fit within the portion of memory allocated to the GPU, with room for the KV cache and system. On a 128 GB machine, the GPU allocation adjustable in the BIOS commonly reaches 96 GB under Windows, and more under Linux through GTT allocation (exact value to be confirmed depending on the motherboard and kernel).

Which catalog models fit on a Ryzen AI Max+ 395 with 128 GB

With approximately 96 GB available for the GPU, a Ryzen AI Max+ 395 workstation can run several models ranging from 118 to 142 billion parameters in Q4 quantization. Here are the models in the quelllm.fr catalog, with indexed Q4 weights and estimates for the other quantizations (Q5 ≈ ×1.2, Q8 ≈ ×1.9, FP16 ≈ ×3.4, estimated).

Beyond that, Step 3.5 Flash (Q4 ~118 GB) and MiniMax-M2.7 (Q4 ~138 GB) exceed the memory of a 128 GB AMD AI PC, even under Linux. They are not realistic options on this platform without more aggressive quantization (Q2 or Q3, quality to be confirmed).

Expected throughput: bandwidth decides, not core count

On a unified-memory chip, token generation is limited by memory bandwidth. Each token requires the GPU to reread all active parameters. The order-of-magnitude calculation is therefore: maximum throughput ≈ bandwidth ÷ bytes read per token.

Prompt processing (prefill), on the other hand, depends on raw compute rather than bandwidth. A context of 30,000 tokens can take several dozen seconds on a Radeon 8060S. Account for this in RAG or long-document analysis use cases.

What about the XDNA 2 NPU? Useful, but not for large models

The NPU integrated into Ryzen AI advertises approximately 50 TOPS in INT8. You need to be precise about what it actually does for a local LLM.

For the models with 118 billion parameters and more listed above, the NPU does not come into play. The Radeon GPU and unified memory do all the work. Our article what is an NPU explores this point further.

Inference tools compatible with Ryzen AI

There are three ways to run a model from the catalog on an AMD AI PC.

In all cases, set the GPU memory to the maximum in the BIOS (an option often called “UMA frame buffer” or “variable graphics memory”), then verify with the runtime that the model is fully loaded on the GPU and not partially on the CPU.

Licenses and benchmarks: what to check before choosing

The license determines professional use. The cited models fall into the following categories.

For benchmarks, quelllm.fr has not measured these models on Ryzen AI. The scores published by the vendors (MMLU for general knowledge, HumanEval and SWE-bench for code, AIME for mathematical reasoning) must be confirmed on the model pages and on theOpen LLM Leaderboard. A benchmark score does not replace testing on your own French documents. Our guide understand the leaderboards explains the pitfalls of reading.

FAQ

Q: Which LLM should I choose for an AMD AI PC with 32 GB of RAM?

None of the models with 118 billion parameters or more in this selection fit in 32 GB. On a Ryzen AI 9 HX laptop, target 7- to 14-billion-parameter models in Q4, which occupy 5 to 10 GB. Our 32 GB configuration guide and the page best 8 GB VRAM LLM list suitable candidates.

Q: Is the Ryzen AI Max+ 395 with 128 GB better than a dedicated graphics card?

It depends on the model you are targeting. A RX 9070 XT with 16 GB offers much higher bandwidth and will be faster on a 30-billion-parameter model. But it will never load a 120-billion-parameter model in Q4. The Ryzen AI Max+ 395 prioritizes capacity over throughput. See our page best LLM for RX 9070 XT to compare.

Q: Can the NPU be used to accelerate Ollama or llama.cpp?

No, not at this time. These two runtimes use the Radeon GPU through Vulkan or ROCm. The XDNA 2 NPU is used only for ONNX models run by the Ryzen AI software stack, for sizes of approximately 1 to 8 billion parameters. For large models, rely solely on the GPU and unified memory.

Q: Do you need Windows or Linux on an AMD AI PC for local inference?

Both work. Windows provides Vulkan through llama.cpp with no special configuration. Linux adds ROCm and lets you allocate more memory to the GPU beyond the BIOS limit, which matters on a 128 GB machine. To load Mixtral 8x22B or dots.llm1 at more than 80 GB, Linux is the safest option.

Q: Which quantization should you choose on a Ryzen AI Max+ 395?

Q4_K_M is the right compromise for models with 118 to 128 billion parameters, around 72 GB. Q5 remains possible for smaller ones like Laguna S 2.1, estimated at around 82 GB. Q8 exceeds the available memory for this entire selection. Keep at least 10 GB free for the KV cache if you use long contexts.

Q: Are these models good in French?

Mistral Small 4 and Mistral Medium 3.5 are trained by a French vendor and perform well in our language. Qwen 3.5 122B-A10B is multilingual and handles French correctly. Precise per-language scores are still to be confirmed on the model pages. Our page best LLM in French compare several models on this criterion.

Conclusion

The best LLM for an AMD AI PC depends first on the available memory. On a Ryzen AI Max+ 395 with 128 GB, Qwen 3.5 122B-A10B, Mistral Small 4, and Nemotron 3 Super 120B fit in Q4 with interactive throughput for MoE architectures. On a Ryzen AI 9 HX with 32 GB, opt for much smaller models. The NPU remains useful for small tasks. To fine-tune the choice for your exact machine, use the configurator or browse the catalog filtered by VRAM.

The hardware for running an LLM locally

To run these models comfortably locally, a RTX 5070 Ti offers an excellent price/performance ratio. Compare prices:

Amazon GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395) →

Affiliate links — QuelLLM may earn a commission on purchases, at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.

Article published and updated on by Mohamed Meguedmi · Data source: /api/models.json · Content license: CC BY 4.0.

An error or update to report? Contribute.