NPU: what is it, and is it useful for a local LLM? ?
A NPU (Neural Processing Unit) is a low-power processor built into the chips of most laptops sold since 2024, dedicated to small, always-on AI tasks such as background blur, live captions, and photo search. Its headline metric, TOPS, says almost nothing about its ability to run a local LLM: generation speed depends on memory bandwidth, and the NPU shares the same slow RAM as the processor.
An NPU, short for Neural Processing Unit, is a processor specialized in neural-network computations, integrated into the chip of most laptops sold since 2024. It continuously runs small AI tasks, such as background blur or live captions, using a few watts. Its headline figure, TOPS, is used as a selling point on every “AI PC” label. As of September 20, 2026, however, it says almost nothing about a machine's ability to run an LLM locally. Here's why, and what to look at instead.
Choosing a machine? Our picks by budget →
Good value for local AI: a GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395).
A mini PC is a complete machine: check the available memory and engine compatibility. It does not replace macOS/MLX or CUDA.
Why this choice? Our complete guide on GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395) →
Compare all options by budget, from €800 to €3,500 →
Small budget: RTX 5060 · Large models: RTX 5090 · Mac Studio.
On the go: which laptop for local AI →
Affiliate links — possible commission at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.
#What is an NPU?
From a computational perspective, a neural network boils down to immense series of matrix multiplications on low-precision numbers, repeated billions of times. A conventional processor can do this, slowly, a handful of operations at a time. A graphics card does it quickly, in parallel across thousands of units, consuming tens to hundreds of watts. An NPU is a piece of silicon designed to do only that, and nothing else: fixed-function compute units for 8- and 4-bit integers, small memories placed right next to them to avoid costly round trips, and a scheduler that moves data as little as possible. The result of this extreme specialization is modest peak speed compared with a GPU, but much higher efficiency per watt, which matters enormously on battery power.
The idea is older than the marketing label attached to it today. Apple has included a Neural Engine in its iPhones since 2017 and in all its Apple Silicon Macs since their launch; Google's Pixel chips have included a TPU block for several generations. What really changed in 2024 is that Intel, AMD, and Qualcomm all placed a substantial NPU in their mainstream laptop processors at the same time, and Microsoft immediately tied an entire range of Windows features to it under the commercial name Copilot+ PC. This convergence, more than the technology itself, explains the flood of “AI PC” labels on store shelves over the past two years.
#CPU, GPU, NPU: the differences
| CPU | GPU | NPU | |
|---|---|---|---|
| Designed for | The general logic, low latency | Massively parallel computing, graphics | Neural network inference only |
| Power consumption under AI load | 15 à 125 W | 30 W (integrated) to 575 W (RTX 5090) | 1 à 10 W |
| Memory used | System RAM | Its own VRAM, or RAM if it is integrated | System RAM |
| Software support | Universal | Very broad: CUDA, ROCm, Metal, Vulkan | Narrow, specific to each manufacturer |
| Playground | Orchestration, with small models as a fallback | LLMs, image generation, training | Lightweight, always-on AI on battery power |
The three chips complement each other; they are not competitors: each exists because it makes a different trade-off between speed, power consumption, and versatility. On an “AI PC,” the NPU handles webcam effects for four hours of video conferencing while the GPU remains idle, saving a substantial portion of the battery life that continuous GPU use would consume. Ask the same machine to run an assistant with 14 billion parameters, and the workload moves entirely to the GPU because that is where both the necessary memory bandwidth and mature software support are available: no mainstream runtime can currently run an LLM of this size on an NPU.
#TOPS: what they measure and what they hide
TOPS stands for “trillion operations per second” (Tera Operations Per Second), almost always measured using 8-bit integers to remain comparable across datasheets. It represents peak compute performance, measured in the lab under favorable conditions. It says nothing about how quickly data reaches those compute units from memory, which is precisely what limits a language model in practice: a data-starved NPU spends most of its time waiting, regardless of how high its TOPS are.
| Chip family | Advertised NPU | Copilot+ eligible (40 TOPS and above) |
|---|---|---|
| Intel Core Ultra Series 1 (Meteor Lake) | ≈ 11 TOPS | No |
| Apple M4 (Neural Engine) | 38 TOPS | Not applicable (macOS) |
| Qualcomm Snapdragon X Elite and X Plus | 45 TOPS | Yes |
| Intel Core Ultra 200V (Lunar Lake) | 48 TOPS | Yes |
| AMD Ryzen AI 300 | 50 TOPS | Yes |
| AMD Ryzen AI Max+ 395 (Strix Halo) | 50 TOPS | Yes |
| Intel Core Ultra 300 (Panther Lake, CES January 2026) | Up to 50 TOPS (NPU 5) | Yes |
| AMD Ryzen AI 400 / PRO 400 “Gorgon Point” (CES January 2026) | Up to 60 TOPS | Yes |
| Qualcomm Snapdragon X2 Elite (CES January 2026) | 80 TOPS | Yes |
This new generation, presented at CES in January 2026 and already available in laptops sold as of September 28, 2026, pushes the displayed figure well beyond the minimum Copilot+ threshold: the Snapdragon X2 Elite nearly doubles its predecessor’s score, reaching 80 TOPS versus 45 previously. None of this changes the conclusion of this page: the point that follows, about shared memory bandwidth, applies identically to these new chips, however impressive their TOPS figure may look on the label.
#Why an NPU Doesn't Accelerate Your LLMs
To produce a single token, a model must reread essentially all of its weights in memory, from the first to the last active parameter. A 14-billion-parameter model in Q4 weighs about 9 GB, so generating 20 tokens per second means moving 180 GB per second between memory and the compute units. The arithmetic itself—the multiplication proper—is cheap compared with this constant transfer. Generation speed therefore follows available memory bandwidth, not the TOPS or TFLOPS compute power advertised on the product sheet, which explains why two chips with very different TOPS ratings can write text at exactly the same speed.
| Where the model is located | Bandwidth | Ceiling for a 9 GB model |
|---|---|---|
| Laptop LPDDR5X (shared by the NPU and integrated GPU) | ≈ 70 to 135 GB/s | ≈ 8 to 15 tok/s |
| Unified memory of the Ryzen AI Max+ 395 | 256 GB/s | ≈ 28 tok/s |
| RTX 4070 (GDDR6X) | 504 GB/s | ≈ 56 tok/s |
| RTX 5070 Ti (GDDR7) | 896 GB/s | ≈ 100 tok/s |
| RTX 5090 (GDDR7) | 1,792 GB/s | ≈ 200 tok/s |
A 50-TOPS NPU, like an 80-TOPS NPU in the very latest chips, sits at the top of the table. It draws on the same shared LPDDR5X as the processor: regardless of its advertised TOPS count, it physically cannot go faster than that memory allows. So a 45-TOPS laptop and an 80-TOPS laptop will write text at exactly the same speed if their memory capacity and bandwidth are identical.
Where an NPU can really help is prompt processing, a phase that is primarily compute-bound rather than bandwidth-bound, because the model processes all the text already present at once instead of generating it word by word. Hybrid configurations that assign this processing to the NPU and generation to the integrated GPU reduce the wait before the first displayed word and overall power consumption on small, well-supported models. This improves battery life and perceived responsiveness, measurable in latency; it is not a way to run larger models or generate continuously faster.
#The real bottleneck: software
- Ollama, llama.cpp, LM Studio
- They run on the GPU or CPU. On a Copilot+ laptop, they use the integrated GPU or processor and ignore the NPU.
- Windows
- Microsoft's embedded models and its Foundry Local environment target the NPU through ONNX Runtime, with a selection of small models.
- AMD
- Ryzen AI Software and the Lemonade server run some ONNX models in hybrid mode, using the NPU and integrated GPU.
- Intel
- OpenVINO can target the NPU for supported models.
- Qualcomm
- The QNN and AI Hub toolchains provide models already converted for the Hexagon NPU.
- Apple
- Core ML uses the Neural Engine, but the LLM tools actually used on Mac (MLX, llama.cpp, Ollama) run on the GPU through Metal.
The common thread across all these tools is that using the NPU means choosing from a short list of models already converted and validated by the manufacturer, generally with 1 to 8 billion parameters, rather than downloading any GGUF found on Hugging Face and running it directly. This conversion requires a different format and toolchain for each manufacturer, which explains why no community project has yet unified NPU support the way llama.cpp did for GPUs.
#What to look at instead before buying
- 01Memory available to the GPUThe VRAM of a dedicated card, or the unified memory of a Mac or a Strix Halo machine. This is what determines which models can load.
- 02The bandwidth of this memoryIt directly determines the number of tokens generated per second.
- 03The software pathNVIDIA with no friction, Apple Silicon nearly as good, AMD properly supported on Linux.
- 04NPU TOPSUseful if you want Windows Copilot+ features and long battery life during video calls. No effect on open-weight LLMs today.
- Hardware for local AI: our machine guides
- Ryzen AI Max+ 395 mini PC: when unified memory changes the game
- Which GPU for a local LLM?
- Microsoft: Copilot+ devices and their NPU
- Source: Wikipedia, history and architecture of NPUs
- Source: AMD press release, Ryzen AI 400 (CES 2026)
#FAQ
What does NPU mean?+
Is an NPU better than a GPU for AI?+
Do Ollama or LM Studio use the NPU?+
Do you need an NPU to run AI locally?+
How many TOPS do you need?+
Do the new 2026 NPU chips change anything for local AI?+
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.