Intermediate 11 minNPU

Snapdragon X Elite: running a local LLM on a PC ARM

Snapdragon X Elite-based Copilot+ PCs have brought ARM architecture back to the center of the Windows world, with an appealing argument for AI: an integrated NPU and several hours of battery life. But running a Snapdragon X Elite LLM locally is not the same experience as on an x86 PC with GPU NVIDIA. This guide separates what really works today from what is still marketing promise, and explains how to get a usable local assistant on these ARM machines.

Choosing a machine? Our picks by budget →

By Thomas P.·Update 2026-08-27·Tested on Windows 11
Recommended hardware

Good value for local AI: a GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395).

A mini PC is a complete machine: check the available memory and engine compatibility. It does not replace macOS/MLX or CUDA.

Why this choice? Our complete guide on GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395) →

Compare all options by budget, from €800 to €3,500 →

Small budget: RTX 5060 · Large models: RTX 5090 · Mac Studio.

On the go: which laptop for local AI →

Affiliate links — possible commission at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.

#Why the Snapdragon X Elite is attracting interest for a local LLM

The Snapdragon X Elite is an ARM SoC from Qualcomm: up to 12 Oryon cores, an integrated Adreno GPU, a Hexagon NPU rated at 45 TOPS, and, most importantly, unified LPDDR5X memory shared by the CPU, GPU, and NPU. This unified memory echoes the Apple Silicon approach: the loaded model is accessible to all compute engines without copying, and you are not limited by separate 8 or 12 GB of VRAM.

The other advantage is the thermal profile. These chips were designed for fanless or nearly silent ultraportables, with battery life measured in tens of hours for office use. For a local LLM running in the background—summarizing emails, rephrasing text, or handling small scripts—this is unprecedented territory: no fan spinning up, no battery draining in twenty minutes.

i
What this guide aims to do
A comfortable local conversational assistant (3B to 24B) on a Windows ARM PC, not a replacement for a RTX 4090. The goal is the best balance of quality, speed, and battery life on this specific hardware.

#The real challenge: Windows on ARM

The first obstacle is not compute power; it’s software compatibility. Windows 11 on ARM runs x86/x64 applications through an emulation layer (Prism), but this emulation costs performance and, for AI, often prevents hardware acceleration from being used. A binary compiled for x86 with AVX instructions will see neither the Adreno GPU nor the Hexagon NPU.

The good news: the main local AI tools now have native ARM64 builds. This is the crucial distinction to keep in mind throughout this guide—a tool that “installs” is not necessarily a “native ARM” tool. Under emulation, you lose much of the machine's benefit.

Ollama
Provides a native Windows ARM64 build. Inference runs on the Oryon CPU and listens by default on http://localhost:11434, as on any other platform.
LM Studio
Provides a native ARM64 version for Windows on Snapdragon, with an ARM-compiled llama.cpp runtime.
llama.cpp
It compiles natively for ARM64 and takes advantage of vector instructions; it's the most optimized underlying engine for these chips.
Avoid
Any x64 version launched under Prism emulation: it runs, but slowly and without access to specialized hardware.
!
Always check the architecture
On the download page, explicitly look for “ARM64” or “Windows on ARM.” If the only option is “Windows x64,” the application will run under emulation and you will lose the machine’s main advantage.

#The Hexagon NPU: myth and reality

This is the heart of the misunderstanding. Copilot+ marketing highlights the Hexagon NPU's 45 TOPS as if every local LLM would automatically benefit from it. In practice, in 2026, almost all LLMs you run through Ollama or LM Studio run on the CPU cores, not the NPU.

Why? Because using the NPU requires models converted and quantized specifically for Qualcomm’s runtime (through QNN — Qualcomm AI Engine Direct — and the ONNX format, generally via ONNX Runtime with the QNN execution provider). These are not the same GGUF files consumed by llama.cpp. The NPU excels at scheduled, predictable workloads; autoregressive text generation, which is dominated by memory bandwidth, is less suitable for it than you might think.

What uses the NPU
Demos and applications packaged through the Qualcomm SDK / AI Hub with pre-converted ONNX models (often small models or vision/audio tasks).
What doesn't use the NPU
Ollama, LM Studio, and llama.cpp in standard use rely on the CPU (and sometimes the Adreno GPU through experimental backends).
Performance in practice
On these machines, generation speed is determined primarily by the Oryon cores and LPDDR5X bandwidth, not the NPU's advertised TOPS.
→
Do not choose this machine for the NPU
If your goal is a fast local chatbot, think of it like any good ARM CPU with fast memory. The NPU is a bonus for specific applications, not the engine of your everyday Ollama.

#Prerequisites and checks

Before installing anything, confirm that you are on an ARM machine and check how much RAM you have—it determines the size of the models you can run, since memory is unified.

PowerShell — check architecture and RAM
# Architecture du processeur (doit renvoyer ARM64)
$env:PROCESSOR_ARCHITECTURE

# Mémoire physique totale, en Go
(Get-CimInstance Win32_ComputerSystem).TotalPhysicalMemory / 1GB
OS
Windows 11 24H2 or later, up to date (recent builds significantly improve Prism emulation and ARM support).
RAM
16 GB minimum for comfortable 3B–9B use; 32 GB to target 20B to 27B models while keeping system headroom.
Storage
Plan for 10 to 30 GB of free space depending on the number of downloaded models (a 7B Q4_K_M weighs ~5 GB).
Realistic expectations
Expect a smooth experience with a small model, a decent one with 8–9B, and a patient wait beyond that. These are not dedicated-GPU throughput levels.

#Install a local LLM step by step

The simplest and most reliable approach today is Ollama in a native ARM64 build, optionally supplemented by a graphical interface. Here's how to do it.

  1. 01
    Download Ollama ARM64
    Go to ollama.com/download and download the Windows installer. In recent versions, the installer detects the ARM64 architecture and installs the native binary. Afterwards, verify that the service is running properly.
  2. 02
    Check the service
    Open a terminal and run « ollama --version ». The daemon listens on http://localhost:11434; a request to this URL should respond with « Ollama is running ».
  3. 03
    Download a first model
    Start small to validate the entire pipeline before loading something heavier. A 3B model in Q4_K_M is ideal for testing responsiveness.
  4. 04
    Chat on the command line
    Launch the model and check the generation speed. If it runs smoothly, gradually increase the size (7B then 14B) as long as the RAM keeps up.
  5. 05
    Add an interface (optional)
    Install LM Studio in the ARM64 version for a complete UI, or connect Open WebUI to the Ollama endpoint if you prefer a web interface.
Terminal — first model
# Modèle léger pour valider la chaîne
ollama pull granite4.2:3b
ollama run granite4.2:3b

# Une fois validé, monter en gamme
ollama pull qwen3.5:9b
→
Measure actual speed
In Ollama, run « ollama run <modèle> --verbose » to display tokens/second at the end of the response. This is the honest metric for comparing models on your machine instead of relying on announcements.

#Recommended models for 16–32 GB of RAM

Because memory is unified and shared with the system, do not reason as if all RAM were available to the model. Always reserve 4 to 6 GB for Windows and your applications. The Q4_K_M memory-footprint benchmarks remain the same as on a GPU: a 3B model ≈ 2 GB, a 9B model ≈ 6-7 GB, and a 24B model ≈ 14 GB.

16 GB of RAM
Comfort zone: 3B to 9B models in Q4_K_M. Granite 4.2 3B for responsiveness, Qwen 3.5 9B or Granite 4.2 8B for quality. A gpt-oss 20B (~14 GB) remains usable but leaves little headroom.
32 GB of RAM
You can comfortably run the 20–27B range—gpt-oss 20B or Mistral Small 24B (~14 GB, the latter particularly comfortable in French), up to Qwen 3.8 27B (~18 GB, 262k context)—in Q4_K_M while leaving plenty of room for the system. A dense 32B or larger model is loadable but slow on this CPU; a MoE such as Qwen 3.6 35B-A3B (~23 GB, 3B active) remains noticeably faster.
Quantization
Q4_K_M is the best quality/size/speed compromise here. Move up to Q5_K_M only if quality is the priority and the RAM allows it; avoid FP16, which is unnecessarily heavy for this class of machine.
Ideal use cases
Rewriting, summarization, translation, light coding assistance, offline Q&A—tasks where a good 9B–24B model is more than sufficient.
i
The right quantization reflex
On a machine where speed comes from the CPU and memory bandwidth, a smaller Q4_K_M model will often provide a better experience than a larger model that struggles. Prioritize smoothness.

#Battery life and silence: the real advantages

This is where the Snapdragon X Elite really stands out. While an x86 gaming laptop with an RTX drains its battery within minutes under AI workloads and sends its fans into overdrive, these ARM machines generate text while staying cool and quiet, whether plugged in or not. For mobile use—working on a train, in a café, or in a meeting room—this is a change in kind, not degree.

In raw speed, expect decent but modest throughput: on the order of several dozen tokens per second on a 3B, dropping to a few tokens per second on a dense 24B. That is enough for a smooth conversational experience with a small model, but more laborious for long generations with a large model. First-token latency remains reasonable thanks to the fast Oryon cores.

Silence
Generation is possible without the fans ramping up, often in near-fanless mode on low-TDP models.
Autonomy
Battery-powered inference is viable, whereas a dedicated GPU requires mains power in almost all cases. Ideal for a mobile local assistant.
Thermals
No aggressive throttling during prolonged light LLM use, unlike x86 ultraportables pushed to their limits.
Trade-off
Throughput that tops out well below a dedicated GPU: convenience comes at the cost of tokens per second on large models.

#Limitations compared with x86 and dedicated GPUs

Let’s be clear to avoid disappointment. A Snapdragon X Elite does not rival a PC equipped with a recent NVIDIA card. An entry-level RTX 3060 12 GB will run a 14B much faster, and a RTX 4090 24 GB is in a completely different category. The Snapdragon’s strength is not raw performance, but performance per watt and mobility.

Still a young ecosystem
ARM64 support is advancing quickly but remains less mature than x86. Some tools, extensions, or GPU backends lag behind or are experimental.
No consumer GPU acceleration
The Adreno backend for llama.cpp remains rudimentary; in practice, inference relies mainly on the CPU.
Underutilized NPU
As noted above, the 45 TOPS do not yet benefit common LLM runtimes.
Size limit
Beyond around twenty billion dense parameters, the experience degrades noticeably; very large models remain the domain of dedicated GPUs (a low-active-parameter MoE is the exception that works better).
!
Setting realistic expectations
Buy (or use) a Snapdragon X Elite for its battery life and silent operation with 3B–24B LLMs. If your priority is running a large dense 32B or 70B model quickly, you need a dedicated GPU, not an ARM PC.

#Go further

The Snapdragon X Elite is best understood by comparing it with other NPU-based approaches and mastering the basics of local installation. These guides continue the discussion:

Ryzen AI 9 HX: does the 50-TOPS NPU really help with LLMs?
The x86 counterpart to this NPU debate, with honest tests that corroborate this guide's conclusions.
Install Ollama: Windows, macOS, and Linux
To master the Ollama daemon and its commands, also applicable to ARM64 builds.
Choosing your quantization (Q4, Q5, Q8)
To fine-tune the quality/speed tradeoff, which is decisive on a machine where the CPU does most of the work.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.