Best LLM for AMD AI PCs (Ryzen AI & NPU)
Un AMD AI PC refers to a machine equipped with a Ryzen AI processor—that is, an APU combining Zen 5 cores, an integrated Radeon GPU, and an XDNA 2 NPU. Choosing the best LLM for an AMD AI PC therefore comes down to a specific question: how much unified memory can the machine allocate to the model, and what bandwidth feeds that memory? The answer is not the same for a Ryzen AI 9 HX laptop with 32 GB and a Ryzen AI Max+ 395 workstation with 128 GB. This article covers both machine families, the models in the quelllm.fr catalog that actually fit in their memory, the role of the NPU, compatible inference tools, licenses, and answers the most frequently asked questions.
Two AMD AI PC families, two memory budgets
The term AMD AI PC covers very different machines. You need to distinguish between them before discussing models.
- Ryzen AI 300 (HX 370, HX 375, 9 365) : Radeon 890M GPU, LPDDR5X memory on a 128-bit bus, bandwidth of approximately 120 GB/s (estimated). Consumer laptops come with 16, 32, or sometimes 64 GB shared between the system and GPU. Our Ryzen AI 9 HX guide details this configuration.
- Ryzen AI Max (385, 390, Max+ 395, codenamed Strix Halo) : Radeon 8060S GPU with 40 compute units on the Max+ 395, a 256-bit memory bus, and approximately 256 GB/s of bandwidth (estimated). Mini PCs and workstations are sold with 64 or 128 GB of soldered LPDDR5X. The Ryzen AI Max+ 395 guide covers BIOS and Linux settings.
The practical rule is simple. The quantized model's size must fit within the portion of memory allocated to the GPU, with room for the KV cache and system. On a 128 GB machine, the GPU allocation adjustable in the BIOS commonly reaches 96 GB under Windows, and more under Linux through GTT allocation (exact value to be confirmed depending on the motherboard and kernel).
Which catalog models fit on a Ryzen AI Max+ 395 with 128 GB
With approximately 96 GB available for the GPU, a Ryzen AI Max+ 395 workstation can run several models ranging from 118 to 142 billion parameters in Q4 quantization. Here are the models in the quelllm.fr catalog, with indexed Q4 weights and estimates for the other quantizations (Q5 ≈ ×1.2, Q8 ≈ ×1.9, FP16 ≈ ×3.4, estimated).
- Qwen 3.5 122B-A10B (Alibaba, Apache 2.0): Q4 ~73 GB, Q5 ~88 GB (estimated), Q8 ~139 GB (does not fit), 262,000-token context. MoE architecture with approximately 10 billion active parameters per token, making it the fastest candidate on this list. The weights are published by Alibaba on Hugging Face.
- Qwen3.8 Flash Next 125B-A6B (Qwen, specific open-weights license): Q4 ~72 GB, 256,000-token context. Only 6 billion active parameters, so throughput is still higher than the previous model (estimated), at the cost of a license that must be read carefully before commercial use.
- Mistral Small 4 (Mistral AI, Apache 2.0): Q4 ~72 GB, 256,000-token context. Good performance in French, published by Mistral AI on Hugging Face.
- Mistral Medium 3.5 128B (Mistral AI, Modified MIT): Q4 ~74 GB, 256,000-token context. Positioned above Small 4 in quality; check the model pages for scores.
- Nemotron 3 Super 120B (NVIDIA, NVIDIA Open Model License): Q4 ~72 GB, 128,000-token context. Geared toward reasoning and agents. Weights available from NVIDIA on Hugging Face.
- Laguna S 2.1 (Poolside, OpenMDW 1.1): Q4 ~68 GB, 262,144-token context. The lightest in the selection, specialized for code, published by Poolside on Hugging Face.
- Mixtral 8x22B Instruct (Mistral AI, Apache 2.0): Q4 ~82 GB, 64 000-token context. It still fits in 96 GB but leaves little room for the KV cache. Older model, mainly useful as a reference.
- dots.llm1 Instruct (Rednote, MIT): ~85 GB Q4, 32,768-token context. Upper limit for a 96 GB allocation, short context.
Beyond that, Step 3.5 Flash (Q4 ~118 GB) and MiniMax-M2.7 (Q4 ~138 GB) exceed the memory of a 128 GB AMD AI PC, even under Linux. They are not realistic options on this platform without more aggressive quantization (Q2 or Q3, quality to be confirmed).
Expected throughput: bandwidth decides, not core count
On a unified-memory chip, token generation is limited by memory bandwidth. Each token requires the GPU to reread all active parameters. The order-of-magnitude calculation is therefore: maximum throughput ≈ bandwidth ÷ bytes read per token.
- MoE model with 10 billion active parameters in Q4 (in the case of Qwen 3.5 122B-A10B): about 5 to 6 GB read per token. With the Ryzen AI Max+ 395's 256 GB/s, that gives a theoretical ceiling of 40 to 50 tokens/s, and a real throughput of 20 to 35 tokens/s once KV cache and runtime overhead are included (estimated, not measured by quelllm.fr on this machine).
- MoE model with 6 billion active parameters (Qwen3.8 Flash Next): estimated real-world throughput of 30 to 50 tokens/s, to be confirmed.
- Dense model with 70 to 120 billion parameters in Q4 : 40 to 70 GB read per token, or 3 to 6 tokens/s. It's readable but slow for interactive use. So check the model architecture on its page before downloading it: the “A10B” or “A6B” label indicates an MoE.
- Ryzen AI 9 HX 370 : at approximately 120 GB/s, the same 120-billion-parameter MoE models do not fit in memory. The problem is not throughput but capacity.
Prompt processing (prefill), on the other hand, depends on raw compute rather than bandwidth. A context of 30,000 tokens can take several dozen seconds on a Radeon 8060S. Account for this in RAG or long-document analysis use cases.
What about the XDNA 2 NPU? Useful, but not for large models
The NPU integrated into Ryzen AI advertises approximately 50 TOPS in INT8. You need to be precise about what it actually does for a local LLM.
- What it does : run models compiled in ONNX format through AMD’s Ryzen AI software stack, in hybrid NPU + GPU mode for generation. The supported models are currently small, ranging from 1 to 8 billion parameters (to be confirmed depending on the software stack version).
- What it doesn't do : speed up a GGUF model loaded in llama.cpp or Ollama. These runtimes use the Radeon GPU via Vulkan or ROCm and ignore the NPU.
- Its appeal : offload the GPU and battery for background tasks (transcription, small resident assistant, embeddings) while the GPU serves a larger model.
For the models with 118 billion parameters and more listed above, the NPU does not come into play. The Radeon GPU and unified memory do all the work. Our article what is an NPU explores this point further.
Inference tools compatible with Ryzen AI
There are three ways to run a model from the catalog on an AMD AI PC.
- llama.cpp with Vulkan backend : the most robust approach on both Windows and Linux. The Vulkan backend works on all integrated Radeon GPUs without ROCm installation. Performance gains from the quantization format (Q4_K_M, Q5_K_M, Q8_0) are comparable to those observed on dedicated GPUs.
- Ollama with ROCm : on Linux, ROCm natively uses the Radeon 8060S. The configuration requires a few environment variables, described in our guide Ollama on AMD GPUs with ROCm. On Windows, Ollama relies on Vulkan (support to be confirmed depending on the version).
- vLLM : possible under Linux with ROCm, but geared toward multi-user servers. For a single workstation, llama.cpp remains simpler. The vLLM documentation lists the supported AMD GPUs.
In all cases, set the GPU memory to the maximum in the BIOS (an option often called “UMA frame buffer” or “variable graphics memory”), then verify with the runtime that the model is fully loaded on the GPU and not partially on the CPU.
Licenses and benchmarks: what to check before choosing
The license determines professional use. The cited models fall into the following categories.
- Apache 2.0 : Qwen 3.5 122B-A10B, Mistral Small 4, Mixtral 8x22B. Free commercial use, modification, and redistribution permitted.
- MIT or Modified MIT : dots.llm1 (MIT), Mistral Medium 3.5 (Modified MIT). Read the clauses added for the modified version.
- Publisher licenses : Nemotron 3 Super 120B (NVIDIA Open Model License), Laguna S 2.1 (OpenMDW 1.1), Qwen3.8 Flash Next (specific open-weights license). These texts generally allow commercial use but add attribution or usage conditions. Our guide to AI Act compliance for open-weight models helps document this choice.
For benchmarks, quelllm.fr has not measured these models on Ryzen AI. The scores published by the vendors (MMLU for general knowledge, HumanEval and SWE-bench for code, AIME for mathematical reasoning) must be confirmed on the model pages and on theOpen LLM Leaderboard. A benchmark score does not replace testing on your own French documents. Our guide understand the leaderboards explains the pitfalls of reading.
FAQ
Q: Which LLM should I choose for an AMD AI PC with 32 GB of RAM?
None of the models with 118 billion parameters or more in this selection fit in 32 GB. On a Ryzen AI 9 HX laptop, target 7- to 14-billion-parameter models in Q4, which occupy 5 to 10 GB. Our 32 GB configuration guide and the page best 8 GB VRAM LLM list suitable candidates.
Q: Is the Ryzen AI Max+ 395 with 128 GB better than a dedicated graphics card?
It depends on the model you are targeting. A RX 9070 XT with 16 GB offers much higher bandwidth and will be faster on a 30-billion-parameter model. But it will never load a 120-billion-parameter model in Q4. The Ryzen AI Max+ 395 prioritizes capacity over throughput. See our page best LLM for RX 9070 XT to compare.
Q: Can the NPU be used to accelerate Ollama or llama.cpp?
No, not at this time. These two runtimes use the Radeon GPU through Vulkan or ROCm. The XDNA 2 NPU is used only for ONNX models run by the Ryzen AI software stack, for sizes of approximately 1 to 8 billion parameters. For large models, rely solely on the GPU and unified memory.
Q: Do you need Windows or Linux on an AMD AI PC for local inference?
Both work. Windows provides Vulkan through llama.cpp with no special configuration. Linux adds ROCm and lets you allocate more memory to the GPU beyond the BIOS limit, which matters on a 128 GB machine. To load Mixtral 8x22B or dots.llm1 at more than 80 GB, Linux is the safest option.
Q: Which quantization should you choose on a Ryzen AI Max+ 395?
Q4_K_M is the right compromise for models with 118 to 128 billion parameters, around 72 GB. Q5 remains possible for smaller ones like Laguna S 2.1, estimated at around 82 GB. Q8 exceeds the available memory for this entire selection. Keep at least 10 GB free for the KV cache if you use long contexts.
Q: Are these models good in French?
Mistral Small 4 and Mistral Medium 3.5 are trained by a French vendor and perform well in our language. Qwen 3.5 122B-A10B is multilingual and handles French correctly. Precise per-language scores are still to be confirmed on the model pages. Our page best LLM in French compare several models on this criterion.
Conclusion
The best LLM for an AMD AI PC depends first on the available memory. On a Ryzen AI Max+ 395 with 128 GB, Qwen 3.5 122B-A10B, Mistral Small 4, and Nemotron 3 Super 120B fit in Q4 with interactive throughput for MoE architectures. On a Ryzen AI 9 HX with 32 GB, opt for much smaller models. The NPU remains useful for small tasks. To fine-tune the choice for your exact machine, use the configurator or browse the catalog filtered by VRAM.
The hardware for running an LLM locally
To run these models comfortably locally, a RTX 5070 Ti offers an excellent price/performance ratio. Compare prices:
Affiliate links — QuelLLM may earn a commission on purchases, at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.