Intermediate 15 minEdge

LLM on Raspberry Pi 5: Local AI embedded

Running an LLM on a Raspberry Pi sounded like a joke two years ago. With the Pi 5 8 GB, its quad-core Cortex-A76 BCM2712 SoC at 2.4 GHz, and the wave of small 2-4B models released in 2026, it has become a real use case. This guide installs Ollama on Raspberry Pi, measures tokens/sec on Qwen 3.5 2B, Granite 4.2 3B, and Qwen 3.5 4B, and shows when an embedded LLM makes sense.

By Mohamed Meguedmi·Update 2026-08-27·Tested on Windows, macOS, and Linux

#Why run an LLM on a Raspberry Pi?

A Raspberry Pi 5 will not replace your RTX 4090. The goal is not raw speed, but embedded use: an €80 board that consumes 5–12 W, fits in a credit-card-sized case, and runs 24 hours a day without noise. For edge use cases—local voice assistant, smart weather station, offline kiosk, mobile robot—it’s sufficient and incomparably cheaper than a mini PC.

The other benefit is absolute privacy. No data leaves the local network. You can set up a family assistant in a child's room without Cloud, telemetry, or a subscription.

i
In two words
Pi 5 + Ollama + a 2-4B Q4-quantized model = a local assistant that responds at 2 to 4 tokens/second. Slow, but autonomous, quiet, and with marginal hardware cost.

#Hardware requirements

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates
Raspberry Pi 5 (8 GB required)
The 4 GB version can run Qwen 3.5 2B but hits its limit beyond that. 8 GB is the minimum for a comfortable experience. The 16 GB version (released in late 2024) makes Qwen 3.5 4B possible without swapping.
A2 microSD card or NVMe SSD
An A1 SD card dramatically slows the model’s initial loading. Ideally, use an NVMe SSD through the official M.2 HAT—500+ MB/s bandwidth, loading Qwen 3.5 4B in 4 s versus 30 s from SD.
Official 27 W power supply
The Pi 5 draws 9–12 W under inference. The official 5 V/5 A USB-C power supply is required to prevent throttling. A generic 15 W power supply under-volts the board.
Active cooling
Essential. The SoC throttles as soon as it reaches 80 °C. The official fan (€5) or an Argon ONE V3 case is sufficient. Without a heatsink, you lose 30–40% throughput after 2 minutes of inference.
Pi OS 64-bit (Bookworm)
Raspberry Pi OS 64-bit based on Debian 12. Ollama does not support 32-bit. Check with uname -m → it should return aarch64.
!
No usable GPU
The Pi 5's VideoCore VII has no official support for LLM inference (no stable Vulkan backend, no mature ML drivers). Everything runs on the 4 ARM CPU cores. This is the main limitation.

#1. Install Ollama on Pi 5

Ollama has officially supported linux/arm64 since 2024. The universal installation script works directly on Pi OS, with no manual compilation required.

Installation Ollama (1 line)
curl -fsSL https://ollama.com/install.sh | sh

The script downloads the ARM64 binary (~150 MB), installs a systemd service, and starts the daemon on port 11434. Allow 1 to 2 minutes.

Check that the service is running
systemctl status ollama
ollama --version
→
Access from the local network
By default, Ollama listens on 127.0.0.1:11434. To use it from another device on the LAN (phone, laptop), edit /etc/systemd/system/ollama.service and add Environment="OLLAMA_HOST=0.0.0.0:11434". Then run sudo systemctl daemon-reload && sudo systemctl restart ollama.

#2. Recommended models for Raspberry Pi

On a Pi 5 8 GB, you have about 6.5 GB of RAM that is actually usable (the system and its cache take the rest). Any LLM exceeding 4–5 GB in Q4 quantization is unusable—it will swap to the microSD card and drop to 0.1 tok/s.

Three models deliver the best results on Pi 5 in August 2026:

Qwen 3.5 2B (Alibaba, Apache 2.0)
The performance/RAM champion. 1.9 GB in Q4_K_M, multimodal, 256k context, excellent for short chats and French. Ideal for a simple assistant or voice control.
Granite 4.2 3B (IBM, Apache 2.0)
Very lean and token-efficient. 2.2 GB in Q4_K_M. A step up in reasoning quality, built to run H24 without saturating RAM. Strong multilingual performance.
Qwen 3.5 4B (Alibaba, Apache 2.0)
The new “small default model” and the most capable of the three. 3.4 GB in Q4_K_M. Fits in RAM on an 8 GB Pi 5, but leaves little headroom—avoid running a heavy GUI alongside it.
Download the three models
ollama pull qwen3.5:2b
ollama pull granite4.2:3b
ollama pull qwen3.5:4b
i
Why not gpt-oss 20B or Mistral Small 24B?
These models weigh ~14 GB in Q4, so they do not fit in the 6.5 GB actually usable on an 8 GB Pi 5. Even a 16 GB Pi would swap them at 0.3–0.5 tok/s—unusable. Stay below 4B on a Pi; these models want a Mac or a GPU.

#3. Tokens/sec benchmarks (orders of magnitude)

Orders of magnitude measured on a Raspberry Pi 5 8 GB, with the official active cooler, an NVMe SSD via an M.2 HAT, 64-bit Pi OS Bookworm, recent Ollama, Q4_K_M quantization, and a 200-token generation prompt. Figures vary by Ollama version and thermals—treat them as rough estimates, not absolute values.

Qwen 3.5 2B Q4_K_M
≈ 4.4 tokens/sec generation · 6.6 tok/s prompt eval · 1.9 GB RAM · 9.6 W
Granite 4.2 3B Q4_K_M
≈ 3.4 tokens/sec generation · 5.1 tok/s prompt eval · 2.2 GB RAM · 10.2 W
Qwen 3.5 4B Q4_K_M
≈ 2.3 tokens/sec generation · 3.6 tok/s prompt eval · 3.4 GB RAM · 11.2 W
Gemma 4 E2B (multimodal reference)
≈ 3.9 tokens/sec · 4.3 GB RAM · 10.4 W. Fast MoE, but more RAM-hungry.
Granite 4.2 3B Q3_K_S (efficient variant)
≈ 3.6 tokens/sec · 1.7 GB RAM · 9.9 W. The lightest of the bunch, ideal when RAM is limited (4 GB Pi).
→
Read these figures
At 5 tok/s, an 80-word response arrives in 20-25 seconds. That’s slower than a cloud assistant but more than sufficient for non-interactive use cases (notifications, batch summaries, asynchronous voice control).
Reproduce these benchmarks locally
ollama run qwen3.5:2b --verbose "Explique en 100 mots ce qu'est un Raspberry Pi."

The --verbose flag displays the following at the end of generation: eval rate (tok/s), prompt eval rate, and total duration. This is the official metric for comparison.

#4. Relevant offline use cases

At 2-5 tok/s, Pi inference is not intended for real-time chat. It shines when latency is not critical, or when privacy and autonomy matter more than speed.

Local voice assistant
Pi 5 + USB microphone + Whisper.cpp tiny + Qwen 3.5 2B + Piper TTS. Question → transcription → response → synthesized voice, all in 8-12 s without the Cloud. Offline home automation, children’s kiosk, educational robot.
Batch summarization of emails or news
Cron job that pulls RSS feeds every hour, summarizes them with Qwen 3.5 4B, and sends them to Telegram/Matrix. Inference latency is invisible when it runs asynchronously.
Log or ticket classification
Edge Pi that tags incoming logs ('critical error', 'spam', 'normal') before sending them to a central server. Saves WAN bandwidth.
Embedded / portable demo
Client presentation, school workshop, conference: Pi 5 in your pocket, USB-C battery, 100% self-contained local AI demo.
Guardrail for other LLMs
Small local model that filters/rephrases a prompt before sending it to a cloud LLM. Useful for anonymizing sensitive data at the edge.
!
When the Pi isn't enough
RAG over >50 documents, long code generation, agents with complex tool calling, step-by-step reasoning (chain-of-thought models such as Qwen 3.8 27B)—for these use cases, choose an M4 Mac mini or a GMKtec/Beelink mini PC instead. The performance-per-euro ratio is much better.

#5. Pi-specific optimizations

#Reduce the context window

By default, Ollama uses num_ctx=2048. On Pi, that's already a lot—each context token consumes RAM and slows pre-evaluation. For a short chat assistant, lower it to 1024.

At runtime in the session
/set parameter num_ctx 1024

#Limit threads

Ollama uses all cores by default. On Pi 5 (4 cores), this is optimal for raw speed but saturates the SoC, triggering thermal throttling after 1–2 minutes. For long-running use, limiting it to 3 cores maintains a stable throughput for longer.

Limit threads via env
OLLAMA_NUM_PARALLEL=1 OLLAMA_NUM_THREAD=3 ollama serve

#Prefer NVMe over microSD

The initial loading of a model from SD reads several GB sequentially. SD A1 → 30–40 s for Phi-3.5 Mini. NVMe → 3–5 s. Once in RAM, inference is identical.

#Drop down one level if you’re short on RAM

On a Pi 5 8 GB with the graphical interface active, Qwen 3.5 4B in Q4_K_M may start swapping. Two options: downgrade to Granite 4.2 3B (2.2 GB) or Qwen 3.5 2B (1.9 GB), freeing 1 to 1.5 GB of RAM without a collapse in quality.

Lighter model
ollama pull granite4.2:3b
→
Recommended headless mode
If you turn the Pi into a dedicated inference server, disable the LXDE desktop: sudo raspi-config → Boot Options → Console. You’ll recover 300-500 MB of RAM—enough room to move comfortably from 2B to 4B.

#Troubleshooting

ollama: command not found after installation
The systemd script failed (often because Pi OS is too old). Check sudo systemctl status ollama. If you see an 'exec format error', you're on 32-bit—reinstall 64-bit Pi OS.
The model loads, but at 0.3 tok/s
You're swapping. Check free -h: if the 'available' column is below 500 MB during inference, your model is too large. Drop down to Granite 4.2 3B or Qwen 3.5 2B, or switch to Q3.
Thermal throttling after 1 minute
The SoC exceeded 80 °C. Check vcgencmd measure_temp during inference. Solution: official fan or Argon ONE V3. Without active cooling, the Pi 5 cannot run inference continuously.
Out of memory halfway through generation
The Linux OOM killer killed Ollama. Reduce num_ctx to 1024, close the browser, or switch to a smaller model. On a 4 GB Pi, stick with Qwen 3.5 2B, the lightest in the 2026 catalog.
Service Ollama does not start at boot
sudo systemctl enable ollama pour activer l'autostart. Vérifiez ensuite avec systemctl is-enabled ollama.

#Go further

Once Ollama is running on your Pi 5, three natural directions to explore further:

Understand quantization to go further
The guide to choosing your quantization compares Q3, Q4, Q5, and Q8 in detail. Critical on Pi, where every 100 MB of RAM counts.
Integrate into a Python app
The guide on integrating Ollama into Python via the REST API shows how to control your Pi 5 from a Flask/FastAPI script—the foundation of any edge assistant.
Automate offline workflows
The guide automate with n8n and Ollama applies as-is on Pi (n8n runs natively on ARM64). Ideal for a standalone AI home automation hub.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.