Intermediate 13 minNPU

Ryzen AI 9 HX: Is the 50-TOPS NPU really useful for LLMs? ?

The Ryzen AI 9 HX 370 proudly posts 50 TOPS on its XDNA 2 NPU—a figure AMD marketing loves to brandish against Intel and Apple. The real question for anyone who wants to run a local LLM is whether this NPU is useful or whether you still have to rely on the integrated Radeon 890M GPU. This guide honestly tests the Ryzen AI stack for LLMs (LM Studio Ryzen AI, Lemonade SDK), compares it with the CPU alone and the iGPU, and gives the verdict—without embellishment.

Choosing a machine? Our picks by budget →

By Mohamed Meguedmi·Update 2026-08-27·Tested on Windows, macOS, and Linux
Recommended hardware

Good value for local AI: a GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395).

A mini PC is a complete machine: check the available memory and engine compatibility. It does not replace macOS/MLX or CUDA.

Why this choice? Our complete guide on GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395) →

Compare all options by budget, from €800 to €3,500 →

Small budget: RTX 5060 · Large models: RTX 5090 · Mac Studio.

On the go: which laptop for local AI →

Affiliate links — possible commission at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.

#Why use an NPU for a local LLM?

An NPU (Neural Processing Unit) is an accelerator dedicated to the INT8 and INT4 matrix operations typical of inference. On the Ryzen AI 9 HX 370, the XDNA 2 block delivers up to 50 TOPS INT8 while consuming about 2 W. At first glance, the tradeoff is unbeatable for portable use—a local LLM that runs without waking the fans or draining the battery in 40 minutes.

The Ryzen AI 9 HX LLM NPU keyword appears everywhere in AMD demos, but the software ecosystem is still young. Most LLM stacks (llama.cpp, Ollama, vLLM) completely ignore the NPU and run everything on the CPU with AVX-512 or on the integrated GPU via Vulkan / ROCm. To use XDNA 2, you need to go through runtimes specifically compiled by AMD or the community.

i
What XDNA 2 can really do
The NPU is designed for INT8/INT4 inference on small models with short contexts. It is not a substitute for a GPU for training, fine-tuning, or running a 32B model. Think “lightweight 3B–8B assistant that responds in the background,” not “AI workstation.”

#XDNA 2 in practice: what the silicon tells us

The XDNA 2 architecture built into Strix Point (Ryzen AI 300 generation) doubles the TOPS compared with XDNA 1 (16 → 50 TOPS) and introduces native support for FP8 and FP16 floating-point formats. It is the first AMD NPU that can handle an LLM decently, but with two major constraints.

Memory bandwidth
The NPU shares system memory (LPDDR5X-7500 or 8000 on Strix Point machines). That's about 120-128 GB/s, shared with the CPU, GPU, and OS. The NPU will never be able to saturate its theoretical compute throughput.
Quantization required
AMD’s official pipeline expects models quantized to INT4 or INT8 through the optimized ONNX format. Standard GGUF Q4_K_M models do not run as-is—you need a conversion.
ONNX precompilation
Each model must be compiled for XDNA 2 (subgraphs are allocated to the NPU, and the rest to the CPU). This is the step that discourages 95% of users.
No multi-user support
The NPU is single-process. If an LLM is using it, no other Windows AI feature (Studio Effects, Recall if you use it) can use it at the same time.

#Hardware and software requirements

CPU
Ryzen AI 9 HX 370, HX 365, or HX PRO 370 (Strix Point, codename Krackan). Ryzen AI 7/5 chips work with a more modest NPU but the same software stack.
RAM
32 GB LPDDR5X minimum for comfortable use. 64 GB if you want to try a 14B model on an iGPU in parallel.
OS
Windows 11 24H2 or later. XDNA NPU drivers for Linux exist (upstream driver since kernel 6.11), but user-space tooling is far behind.
NPU drivers
AMD Ryzen AI Software 1.3 or later (separate download from amd.com, not included in the Adrenalin driver).
Storage
~20 GB free for precompiled ONNX models (larger than equivalent GGUFs).
!
Check the NPU first
In Device Manager → Neural Processors → "AMD Ryzen AI NPU" should appear without a yellow triangle. If not, you have a generic chipset driver: install the dedicated AMD Ryzen AI Software package.

#1. LM Studio Ryzen AI: the simplest path

AMD distributes an official fork of LM Studio preconfigured for XDNA 2. This is the recommended route if you want to test it without touching Python or ONNX Runtime.

  1. 01
    Download the binary
    Go to amd.com/lmstudio (official “LM Studio for Ryzen AI” page). The Windows binary is signed by AMD.
  2. 02
    Standard installation
    The installer places LM Studio in Program Files and automatically configures the ONNX RyzenAI Execution Provider runtime.
  3. 03
    Select an NPU-optimized model
    In the Discover tab, filter by the “Ryzen AI” badge. You will find Qwen 3.5 9B, Granite 4.2 3B, and Granite 4.2 8B precompiled for XDNA 2 (suffix -hybrid).
  4. 04
    Load and test
    On startup, LM Studio displays a runtime selector: GPU, CPU, or “Ryzen AI Hybrid” (NPU + iGPU). Choose Hybrid to use the NPU.
  5. 05
    Verify that the NPU is doing the work
    Open Task Manager → the Performance tab → NPU. The graph should rise to 60–95% during generation. If it stays at 0, the runtime fell back to the CPU.
Check the NPU from the command line
# Lister les périphériques de calcul disponibles
Get-PnpDevice -Class "ComputeAccelerator" | Format-Table FriendlyName, Status

# Doit afficher : AMD Ryzen AI NPU - OK

#2. Lemonade SDK: going beyond LM Studio

Lemonade SDK is AMD’s official tool for developers. It serves two purposes: (1) compile your own GGUF/HF models to the hybrid ONNX format, (2) expose an OpenAI-compatible API usable from Open WebUI, Continue.dev, etc.

Lemonade SDK installation
# Python 3.11 recommandé
python -m venv .venv
.\.venv\Scripts\Activate.ps1

# Installation du SDK
pip install lemonade-sdk[npu]

# Test : lister les modèles déjà disponibles
lemonade list
Launch an OpenAI-compatible server
# Charge Qwen 3.5 9B précompilé en mode hybride NPU+iGPU
lemonade serve --model qwen3.5-9b-instruct-hybrid --port 8000

# L'endpoint http://localhost:8000/v1/chat/completions est désormais utilisable
# par toute app qui parle OpenAI (Open WebUI, Continue, etc.)
→
Compile a custom model
For your own models, the `lemonade onnx-export --source granite-4.2-8b-instruct --target hybrid` command produces a precompiled ONNX directory. Allow 15–30 minutes for compilation the first time, and about 8–12 GB of disk space for an 8B in INT4.

#3. Fall back to the Radeon 890M: often the best choice

The Ryzen AI 9 HX 370 also includes a Radeon 890M iGPU (RDNA 3.5, 16 CU). On LM Studio standard or Ollama, this iGPU can be used through Vulkan and sometimes beats the NPU in raw throughput. Here's how to use it.

Ollama on Radeon 890M (Vulkan)
# Installation Ollama Windows (build avec Vulkan)
winget install Ollama.Ollama

# Forcer le backend Vulkan
$env:OLLAMA_VULKAN = "1"

# Lancer un Qwen 3.5 9B
ollama serve
ollama run qwen3.5:9b
Dynamic VRAM allocation
The Radeon 890M shares system RAM. In the UEFI BIOS, look for “UMA Frame Buffer Size” and set it to at least 8 GB (16 GB if possible) before any LLM testing.
Driver
Adrenalin 24.10 or later is required to benefit from the optimized Vulkan compute path. Lenovo/Asus OEM drivers are often 2 versions behind.
ROCm on Linux
ROCm 6.3+ has officially supported Strix Point since April 2026. On Ubuntu 24.04, it is workable, but the iGPU remains limited by its shared memory bandwidth.

#Benchmarks: NPU vs iGPU vs CPU

Orders of magnitude measured on an Asus Zenbook S 14 (Ryzen AI 9 HX 370, 32 GB LPDDR5X-7500, Windows 11 24H2, AMD Ryzen AI Software 1.3, Adrenalin 24.10.1). Short prompt (~50 tokens), 256-token generation, temperature 0.7. Average of 5 runs.

Qwen 3.5 9B — AVX-512 CPU
4.1 tok/s · 28 W package · audible fan
Qwen 3.5 9B — Radeon 890M (Vulkan)
9.8 tok/s · 35 W package · sustained fan
Qwen 3.5 9B hybrid — NPU + iGPU (LM Studio Ryzen AI)
11.2 tok/s · 18 W package · quiet fan
Granite 4.2 3B — pure NPU
24 tok/s · 9 W package · fan off
Granite 4.2 8B — NPU + iGPU
13.5 tok/s · 19 W package · quiet fan
Gemma 4 12B — Radeon 890M (Vulkan)
4.2 tok/s · 38 W · too large for the NPU alone
i
What these figures show
The NPU is not the fastest in absolute terms—the iGPU can come close on 7B–8B models. But at comparable throughput, the NPU consumes 2× less power and produces 2× less heat. That is its real value: running an LLM session for 3 hours on battery without thermal degradation.

#Use cases where the NPU shines

Always-on assistance chatbot
A Granite 4.2 3B or Qwen 3.5 4B uses ~5 W and remains available H24 without any noticeable battery impact. Ideal for integration into a productivity app.
Automatic email / document summarization
Short batch task, 4–8k-token context, 3B–8B models. Perfect for the NPU: fan off, silent background processing while you work.
Real-time speech-to-text + summarization
Whisper (on the iGPU) + Granite 4.2 8B (on the NPU) in a pipeline — possible precisely because the two blocks are independent.
Extended airplane mode
On battery power, using the NPU + iGPU enables 4–5 hours of active LLM use versus 1 hour 30 minutes on a pure CPU/iGPU setup.
→
Combine the NPU and iGPU intelligently
The “Hybrid” mode of LM Studio Ryzen AI automatically distributes the layers: Attention blocks on the NPU (dense INT4), and FFNs on the iGPU, where memory bandwidth is better utilized. This mode provides the best performance-to-power balance, not the NPU alone.

#Pitfalls and limitations to know

Limited context
Most precompiled AMD models top out at 4096 context tokens. For 32k or more, you have to recompile—and that uses much more memory.
No models > 8B on NPU alone
Beyond 8B, even in INT4, the NPU lacks bandwidth. Beyond 13B, the iGPU necessarily takes over.
Unavailable models
Qwen 3.5, DeepSeek, and Gemma 4 are not all officially precompiled. Check the Ryzen AI hub before buying a configuration expecting to run a specific model.
Breaking Windows update
Some 24H2 cumulative updates temporarily disabled NPU enumeration. Keep the AMD Ryzen AI Software package up to date, and check Device Manager after each cumulative update.
No fine-tuning
XDNA 2 only handles inference. To train or LoRA-tune, you need a dedicated GPU—period.
!
The NPU does not replace a dedicated GPU
If your goal is to run a Qwen 3.6 35B, a 70B model, or fine-tune, the Ryzen AI 9 HX 370 is not the machine for you. A desktop PC with RTX 4070 / 4080 / 4090 remains unbeatable. The NPU is a mobility solution—not a substitute for a GPU.

#Go further

Three paths for going deeper, depending on your next step:

Understanding quantization
The XDNA 2 NPU only handles INT4/INT8. The “Choosing Your Quantization (Q4, Q5, Q8, FP16)” guide explains why Q4_K_M remains the CPU/GPU standard and what INT4 changes on an NPU.
Compare with a real GPU configuration
If you are considering a desktop, the guide “Choosing Your GPU for Local AI” compares RTX 4070, 4090, M-Max, and gives purchasing thresholds by use case.
More conventional Ollama setup
To stay on the standard stack with your Radeon 890M, the installation guide for Ollama on Windows covers Vulkan configuration and shared VRAM/RAM management.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.

Prices in euros (€) are French market prices including VAT, checked by QuelLLM. US prices differ: the Amazon buttons show the current US price.