BestLLMfor Your hardware. Your LLM. Your call.
◆ The kits◆ Kits APIOpen data Find my LLM
Guide · 2026-09-20

What Is an NPU? And Can It Actually Run a Local LLM?

◆ Local AI — Your private ChatGPT, free, on your own machine, in an hour · $24 · or all kits $49 →

Every new laptop advertises an NPU and a TOPS number. What the chip does, how it differs from a GPU, and why it matters less for local LLMs than the sticker suggests.

By Mohamed Meguedmi·Last updated 2026-09-20·9 min read·Tested on Windows, macOS, Linux

Key takeaways

  • An NPU (neural processing unit) is a processor block built to run neural networks at very low power. It sits on the same chip as the CPU and integrated GPU in most laptops sold since 2024.
  • It is designed for small, always-on tasks: background blur, live captions, noise removal, photo search. It does them at a few watts instead of draining the battery on the GPU.
  • TOPS (trillions of operations per second) is the headline spec. Microsoft's Copilot+ PC label requires 40 TOPS or more.
  • For local LLMs, TOPS is the wrong number to shop by. Text generation is limited by memory bandwidth and capacity, and an NPU shares the same system RAM as the CPU.
  • As of September 20, 2026, the mainstream local runtimes (Ollama, llama.cpp, LM Studio) run on the GPU or CPU, not the NPU. NPU paths exist, but through vendor-specific toolchains.

What an NPU is

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • 30-day refund

A neural network is, computationally, a very long sequence of matrix multiplications on low-precision numbers. A CPU can do that, slowly. A GPU does it fast, at tens to hundreds of watts. An NPU is silicon that does nothing else: fixed-function arithmetic units for 8-bit and 4-bit integer math, small on-chip memories placed next to them, and a scheduler that moves data as little as possible. The result is modest peak speed but excellent performance per watt.

The idea is older than the branding (the Wikipedia entry on neural processing units traces the lineage). Apple has shipped a Neural Engine in iPhones since 2017 and in every Apple Silicon Mac; Google's phones have had a TPU block for years. What changed in 2024 is that Intel, AMD and Qualcomm all put sizable NPUs into PC processors, and Microsoft tied a Windows feature tier to them.

NPU vs GPU vs CPU

CPUGPUNPU
Built forGeneral-purpose logic, low latencyMassively parallel math, graphicsNeural-network inference only
Typical power under AI load15–125 W30 W (integrated) to 575 W (RTX 5090)1–10 W
MemorySystem RAMOwn VRAM (discrete) or system RAM (integrated)System RAM
Software supportUniversalVery broad (CUDA, ROCm, Metal, Vulkan)Narrow, vendor-specific
Sweet spotOrchestration, small models in a pinchLLMs, image generation, trainingSustained light AI on battery

The three are complementary, not competing. On an "AI PC" the NPU handles the webcam effects during a four-hour call while the GPU stays asleep. Ask the same machine to run a 14B-parameter chatbot and the work goes to the GPU, because that is where the memory bandwidth and the software are.

What TOPS means, and what it hides

TOPS counts how many trillion integer operations the NPU can perform per second, almost always at INT8 precision. It is a peak compute figure. It says nothing about how fast data reaches those arithmetic units, and that is precisely what limits a language model.

Chip familyNPU ratingCopilot+ eligible (40+ TOPS)
Intel Core Ultra Series 1 (Meteor Lake)≈ 11 TOPSNo
Apple M4 Neural Engine38 TOPSNot applicable (macOS)
Qualcomm Snapdragon X Elite / X Plus45 TOPSYes
Intel Core Ultra 200V (Lunar Lake)48 TOPSYes
AMD Ryzen AI 300 series50 TOPSYes
AMD Ryzen AI Max+ 395 (Strix Halo)50 TOPSYes

Manufacturer figures. Vendors do not measure TOPS identically, so treat differences of a few points as noise. Microsoft documents the Copilot+ requirement and the supported silicon in its NPU developer guide.

Why an NPU does not make local LLMs fast

To produce one token, a model has to read essentially all of its active weights from memory. A 14B model at 4-bit is about 9 GB; generating 20 tokens per second means streaming roughly 180 GB through the processor every second. The arithmetic itself is cheap by comparison. This is why generation speed tracks memory bandwidth, not compute.

Where the model sitsBandwidthRough ceiling for a 9 GB model
Laptop LPDDR5X (what the NPU and iGPU share)≈ 70–135 GB/s≈ 8–15 tok/s
Ryzen AI Max+ 395 unified memory256 GB/s≈ 28 tok/s
RTX 4070 (GDDR6X)504 GB/s≈ 56 tok/s
RTX 5070 Ti (GDDR7)896 GB/s≈ 100 tok/s
RTX 5090 (GDDR7)1,792 GB/s≈ 200 tok/s

Ceiling = bandwidth ÷ model size, an upper bound that real systems approach but do not reach. Model size from the BestLLMfor catalog (Qwen 3 14B, 4-bit).

An NPU rated at 50 TOPS sits on the first row. It draws from the same LPDDR5X as the CPU, so for token generation it cannot outrun that memory, however many TOPS it has. A 45 TOPS laptop and a 50 TOPS laptop will generate text at essentially the same speed if their RAM is the same. The full argument, with the roofline model, is in why VRAM matters more than TFLOPS.

Where an NPU can help an LLM is prompt processing, the compute-heavy phase where the model reads your input. Hybrid schemes that run prefill on the NPU and generation on the integrated GPU cut time-to-first-token and power draw on small models. That is a battery-life win, not a way to run bigger models.

The software gap

The second obstacle is support. CUDA, Metal, ROCm and Vulkan are mature targets that every local runtime speaks. NPUs each need their own compiler stack and model format:

  • Ollama, llama.cpp, LM Studio: run on GPU and CPU. On a Copilot+ laptop they use the integrated GPU or the CPU and ignore the NPU.
  • Windows: Microsoft's on-device models and its Foundry Local runtime target the NPU through ONNX Runtime, with a curated set of small models.
  • AMD: Ryzen AI Software and the Lemonade server run selected ONNX models in hybrid NPU + iGPU mode.
  • Intel: OpenVINO can target the NPU for supported models.
  • Qualcomm: the QNN toolchain and AI Hub provide pre-converted models for the Hexagon NPU.
  • Apple: Core ML uses the Neural Engine, but the LLM tools people actually use on Mac (MLX, llama.cpp, Ollama) run on the GPU through Metal.

The common thread: NPU inference means a short list of vendor-converted models, typically 1B to 8B parameters, rather than "pull any GGUF and run it."

What to look at instead when buying for local AI

  1. Memory capacity the GPU can use. VRAM on a discrete card, or unified memory on a Mac or Strix Halo machine. This decides which models load at all. See what VRAM is and how much you need.
  2. Memory bandwidth. This decides tokens per second.
  3. Software path. NVIDIA is frictionless, Apple Silicon is close, AMD is good on Linux.
  4. NPU TOPS. Relevant if you want Windows Copilot+ features and long battery life during calls. Irrelevant for running open-weight LLMs today.

Concrete options with numbers: the Ryzen AI Max+ 395 mini PC page covers the one x86 platform where unified memory changes the picture, and the AI hardware hub lists the Macs and GPUs we track. For model sizing by memory, use the VRAM calculator. The underlying data is open through the BestLLMfor public API (CC BY 4.0) and our MCP server.

The verdict

An NPU is a genuinely useful chip for what it was built for: keeping light AI features running all day on a few watts. It is not a hidden LLM accelerator. If your goal is running open-weight models locally, buy memory capacity and bandwidth, treat the NPU as a free extra, and ignore TOPS comparisons between otherwise similar laptops. If vendor toolchains mature and mainstream runtimes gain NPU backends, the prefill and efficiency gains will arrive as a software update. The memory ceiling will not move.

Frequently asked questions

What does NPU stand for?

Neural processing unit. It is a processor block dedicated to running neural networks efficiently, integrated into most recent laptop and phone chips alongside the CPU and GPU.

Is an NPU better than a GPU for AI?

It is more power-efficient for small, continuous tasks. A GPU is far faster and far better supported for large models such as LLMs and image generators. They serve different jobs.

Can Ollama or LM Studio use my NPU?

As of September 2026, no. They run on the GPU (CUDA, ROCm, Metal, Vulkan) or the CPU. Using the NPU requires vendor toolchains such as ONNX Runtime with Foundry Local, AMD's Ryzen AI Software, Intel OpenVINO or Qualcomm QNN, each with its own limited model list.

Do I need an NPU to run AI locally?

No. Any machine with a capable GPU or enough unified memory runs local models without one. An NPU is only required for Windows Copilot+ features.

What is a good TOPS number?

40 TOPS is the threshold for Microsoft's Copilot+ PC label, and current laptop chips sit between 45 and 50. Above that threshold, differences have little practical effect today, and none on LLM generation speed.

Does a desktop PC have an NPU?

Usually not a meaningful one. Some recent desktop processors include a small NPU, but a desktop with a discrete graphics card has far more AI compute in the GPU, without the battery constraint the NPU exists to solve.

Recommended hardware

A current option for local AI: GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395). Match memory to your model and software. A mini PC is a complete PC alternative; Mac/MLX and CUDA instructions require compatible hardware.

Amazon Check GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395) price →

As an Amazon Associate, BestLLMfor earns from qualifying purchases, at no extra cost to you. It does not influence our independent rankings.

Did this guide help?

Found an error or have feedback? Let us know — it helps everyone who reads this guide.