Beginner 11 minNPU

NPU: what is it, and is it useful for a local LLM? ?

Direct response

A NPU (Neural Processing Unit) is a low-power processor built into the chips of most laptops sold since 2024, dedicated to small, always-on AI tasks such as background blur, live captions, and photo search. Its headline metric, TOPS, says almost nothing about its ability to run a local LLM: generation speed depends on memory bandwidth, and the NPU shares the same slow RAM as the processor.

An NPU, short for Neural Processing Unit, is a processor specialized in neural-network computations, integrated into the chip of most laptops sold since 2024. It continuously runs small AI tasks, such as background blur or live captions, using a few watts. Its headline figure, TOPS, is used as a selling point on every “AI PC” label. As of September 20, 2026, however, it says almost nothing about a machine's ability to run an LLM locally. Here's why, and what to look at instead.

Choosing a machine? Our picks by budget →

By Mohamed Meguedmi·Update 2026-09-28·Tested on Windows, macOS, and Linux
Recommended hardware

Good value for local AI: a GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395).

A mini PC is a complete machine: check the available memory and engine compatibility. It does not replace macOS/MLX or CUDA.

Why this choice? Our complete guide on GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395) →

Compare all options by budget, from €800 to €3,500 →

Small budget: RTX 5060 · Large models: RTX 5090 · Mac Studio.

On the go: which laptop for local AI →

Affiliate links — possible commission at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.

#What is an NPU?

From a computational perspective, a neural network boils down to immense series of matrix multiplications on low-precision numbers, repeated billions of times. A conventional processor can do this, slowly, a handful of operations at a time. A graphics card does it quickly, in parallel across thousands of units, consuming tens to hundreds of watts. An NPU is a piece of silicon designed to do only that, and nothing else: fixed-function compute units for 8- and 4-bit integers, small memories placed right next to them to avoid costly round trips, and a scheduler that moves data as little as possible. The result of this extreme specialization is modest peak speed compared with a GPU, but much higher efficiency per watt, which matters enormously on battery power.

The idea is older than the marketing label attached to it today. Apple has included a Neural Engine in its iPhones since 2017 and in all its Apple Silicon Macs since their launch; Google's Pixel chips have included a TPU block for several generations. What really changed in 2024 is that Intel, AMD, and Qualcomm all placed a substantial NPU in their mainstream laptop processors at the same time, and Microsoft immediately tied an entire range of Windows features to it under the commercial name Copilot+ PC. This convergence, more than the technology itself, explains the flood of “AI PC” labels on store shelves over the past two years.

#CPU, GPU, NPU: the differences

CPUGPUNPU
Designed forThe general logic, low latencyMassively parallel computing, graphicsNeural network inference only
Power consumption under AI load15 à 125 W30 W (integrated) to 575 W (RTX 5090)1 à 10 W
Memory usedSystem RAMIts own VRAM, or RAM if it is integratedSystem RAM
Software supportUniversalVery broad: CUDA, ROCm, Metal, VulkanNarrow, specific to each manufacturer
PlaygroundOrchestration, with small models as a fallbackLLMs, image generation, trainingLightweight, always-on AI on battery power

The three chips complement each other; they are not competitors: each exists because it makes a different trade-off between speed, power consumption, and versatility. On an “AI PC,” the NPU handles webcam effects for four hours of video conferencing while the GPU remains idle, saving a substantial portion of the battery life that continuous GPU use would consume. Ask the same machine to run an assistant with 14 billion parameters, and the workload moves entirely to the GPU because that is where both the necessary memory bandwidth and mature software support are available: no mainstream runtime can currently run an LLM of this size on an NPU.

#TOPS: what they measure and what they hide

TOPS stands for “trillion operations per second” (Tera Operations Per Second), almost always measured using 8-bit integers to remain comparable across datasheets. It represents peak compute performance, measured in the lab under favorable conditions. It says nothing about how quickly data reaches those compute units from memory, which is precisely what limits a language model in practice: a data-starved NPU spends most of its time waiting, regardless of how high its TOPS are.

Manufacturer figures · measurement methods differ, so a gap of a few points is not significant
Chip familyAdvertised NPUCopilot+ eligible (40 TOPS and above)
Intel Core Ultra Series 1 (Meteor Lake)≈ 11 TOPSNo
Apple M4 (Neural Engine)38 TOPSNot applicable (macOS)
Qualcomm Snapdragon X Elite and X Plus45 TOPSYes
Intel Core Ultra 200V (Lunar Lake)48 TOPSYes
AMD Ryzen AI 30050 TOPSYes
AMD Ryzen AI Max+ 395 (Strix Halo)50 TOPSYes
Intel Core Ultra 300 (Panther Lake, CES January 2026)Up to 50 TOPS (NPU 5)Yes
AMD Ryzen AI 400 / PRO 400 “Gorgon Point” (CES January 2026)Up to 60 TOPSYes
Qualcomm Snapdragon X2 Elite (CES January 2026)80 TOPSYes

This new generation, presented at CES in January 2026 and already available in laptops sold as of September 28, 2026, pushes the displayed figure well beyond the minimum Copilot+ threshold: the Snapdragon X2 Elite nearly doubles its predecessor’s score, reaching 80 TOPS versus 45 previously. None of this changes the conclusion of this page: the point that follows, about shared memory bandwidth, applies identically to these new chips, however impressive their TOPS figure may look on the label.

#Why an NPU Doesn't Accelerate Your LLMs

To produce a single token, a model must reread essentially all of its weights in memory, from the first to the last active parameter. A 14-billion-parameter model in Q4 weighs about 9 GB, so generating 20 tokens per second means moving 180 GB per second between memory and the compute units. The arithmetic itself—the multiplication proper—is cheap compared with this constant transfer. Generation speed therefore follows available memory bandwidth, not the TOPS or TFLOPS compute power advertised on the product sheet, which explains why two chips with very different TOPS ratings can write text at exactly the same speed.

Theoretical ceiling = bandwidth ÷ model size (Qwen 3 14B in Q4, 9 GB in the QuelLLM catalog). Real systems approach it without reaching it.
Where the model is locatedBandwidthCeiling for a 9 GB model
Laptop LPDDR5X (shared by the NPU and integrated GPU)≈ 70 to 135 GB/s≈ 8 to 15 tok/s
Unified memory of the Ryzen AI Max+ 395256 GB/s≈ 28 tok/s
RTX 4070 (GDDR6X)504 GB/s≈ 56 tok/s
RTX 5070 Ti (GDDR7)896 GB/s≈ 100 tok/s
RTX 5090 (GDDR7)1,792 GB/s≈ 200 tok/s

A 50-TOPS NPU, like an 80-TOPS NPU in the very latest chips, sits at the top of the table. It draws on the same shared LPDDR5X as the processor: regardless of its advertised TOPS count, it physically cannot go faster than that memory allows. So a 45-TOPS laptop and an 80-TOPS laptop will write text at exactly the same speed if their memory capacity and bandwidth are identical.

Where an NPU can really help is prompt processing, a phase that is primarily compute-bound rather than bandwidth-bound, because the model processes all the text already present at once instead of generating it word by word. Hybrid configurations that assign this processing to the NPU and generation to the integrated GPU reduce the wait before the first displayed word and overall power consumption on small, well-supported models. This improves battery life and perceived responsiveness, measurable in latency; it is not a way to run larger models or generate continuously faster.

#The real bottleneck: software

Ollama, llama.cpp, LM Studio
They run on the GPU or CPU. On a Copilot+ laptop, they use the integrated GPU or processor and ignore the NPU.
Windows
Microsoft's embedded models and its Foundry Local environment target the NPU through ONNX Runtime, with a selection of small models.
AMD
Ryzen AI Software and the Lemonade server run some ONNX models in hybrid mode, using the NPU and integrated GPU.
Intel
OpenVINO can target the NPU for supported models.
Qualcomm
The QNN and AI Hub toolchains provide models already converted for the Hexagon NPU.
Apple
Core ML uses the Neural Engine, but the LLM tools actually used on Mac (MLX, llama.cpp, Ollama) run on the GPU through Metal.

The common thread across all these tools is that using the NPU means choosing from a short list of models already converted and validated by the manufacturer, generally with 1 to 8 billion parameters, rather than downloading any GGUF found on Hugging Face and running it directly. This conversion requires a different format and toolchain for each manufacturer, which explains why no community project has yet unified NPU support the way llama.cpp did for GPUs.

#What to look at instead before buying

  1. 01
    Memory available to the GPU
    The VRAM of a dedicated card, or the unified memory of a Mac or a Strix Halo machine. This is what determines which models can load.
  2. 02
    The bandwidth of this memory
    It directly determines the number of tokens generated per second.
  3. 03
    The software path
    NVIDIA with no friction, Apple Silicon nearly as good, AMD properly supported on Linux.
  4. 04
    NPU TOPS
    Useful if you want Windows Copilot+ features and long battery life during video calls. No effect on open-weight LLMs today.

#FAQ

What does NPU mean?+
Neural Processing Unit, or neural processing unit. This is a processor block dedicated to efficiently running neural networks, found in most recent laptop and phone chips alongside the CPU and GPU. It continuously handles small AI tasks, such as background blur in video calls, without overheating the battery or quickly draining its charge.
Is an NPU better than a GPU for AI?+
It is significantly more energy-efficient for small continuous workloads, using around 1 to 10 W compared with 30 W or more for a GPU. A GPU remains much faster and much better supported by software for large models, LLMs, and image generators. They are not competitors: they simply do different work, and not with the same model sizes.
Do Ollama or LM Studio use the NPU?+
No, as of September 28, 2026. They run on the GPU via CUDA, ROCm, Metal, or Vulkan depending on the hardware, or on the processor as a last resort. Using the NPU requires each manufacturer's own tools: ONNX Runtime for Windows, Ryzen AI Software for AMD, and OpenVINO for Intel, each with its limited list of already converted models.
Do you need an NPU to run AI locally?+
No, absolutely not. Any machine with a decent GPU or enough unified memory can run local models without an NPU, just as was already the case before 2024. An NPU is required only to enable Windows-specific Copilot+ features, not for Ollama, LM Studio, or any other common inference tool.
How many TOPS do you need?+
40 TOPS is the threshold for Microsoft's Copilot+ PC label. Chips launched at CES in January 2026—Panther Lake, Ryzen AI 400, and Snapdragon X2 Elite—now reach 50, 60, and up to 80 TOPS. Beyond the Copilot+ threshold, the difference has little perceptible practical effect, and absolutely none on the generation speed of a local LLM.
Do the new 2026 NPU chips change anything for local AI?+
No, not for LLMs. Panther Lake, Ryzen AI 400, and Snapdragon X2 Elite increase the displayed TOPS, but the NPU continues to draw from the same system memory as the processor, with no associated bandwidth gain. The bottleneck limiting a language model’s generation speed remains exactly the same as with the previous generation of chips.

Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.

Prices in euros (€) are French market prices including VAT, checked by QuelLLM. US prices differ: the Amazon buttons show the current US price.