LLM on Raspberry Pi 5: Local AI embedded
Running an LLM on a Raspberry Pi sounded like a joke two years ago. With the Pi 5 8 GB, its quad-core Cortex-A76 BCM2712 SoC at 2.4 GHz, and the wave of small 2-4B models released in 2026, it has become a real use case. This guide installs Ollama on Raspberry Pi, measures tokens/sec on Qwen 3.5 2B, Granite 4.2 3B, and Qwen 3.5 4B, and shows when an embedded LLM makes sense.
#Why run an LLM on a Raspberry Pi?
A Raspberry Pi 5 will not replace your RTX 4090. The goal is not raw speed, but embedded use: an €80 board that consumes 5–12 W, fits in a credit-card-sized case, and runs 24 hours a day without noise. For edge use cases—local voice assistant, smart weather station, offline kiosk, mobile robot—it’s sufficient and incomparably cheaper than a mini PC.
The other benefit is absolute privacy. No data leaves the local network. You can set up a family assistant in a child's room without Cloud, telemetry, or a subscription.
#Hardware requirements
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
- Raspberry Pi 5 (8 GB required)
- The 4 GB version can run Qwen 3.5 2B but hits its limit beyond that. 8 GB is the minimum for a comfortable experience. The 16 GB version (released in late 2024) makes Qwen 3.5 4B possible without swapping.
- A2 microSD card or NVMe SSD
- An A1 SD card dramatically slows the model’s initial loading. Ideally, use an NVMe SSD through the official M.2 HAT—500+ MB/s bandwidth, loading Qwen 3.5 4B in 4 s versus 30 s from SD.
- Official 27 W power supply
- The Pi 5 draws 9–12 W under inference. The official 5 V/5 A USB-C power supply is required to prevent throttling. A generic 15 W power supply under-volts the board.
- Active cooling
- Essential. The SoC throttles as soon as it reaches 80 °C. The official fan (€5) or an Argon ONE V3 case is sufficient. Without a heatsink, you lose 30–40% throughput after 2 minutes of inference.
- Pi OS 64-bit (Bookworm)
- Raspberry Pi OS 64-bit based on Debian 12. Ollama does not support 32-bit. Check with uname -m → it should return aarch64.
#1. Install Ollama on Pi 5
Ollama has officially supported linux/arm64 since 2024. The universal installation script works directly on Pi OS, with no manual compilation required.
The script downloads the ARM64 binary (~150 MB), installs a systemd service, and starts the daemon on port 11434. Allow 1 to 2 minutes.
#2. Recommended models for Raspberry Pi
On a Pi 5 8 GB, you have about 6.5 GB of RAM that is actually usable (the system and its cache take the rest). Any LLM exceeding 4–5 GB in Q4 quantization is unusable—it will swap to the microSD card and drop to 0.1 tok/s.
Three models deliver the best results on Pi 5 in August 2026:
- Qwen 3.5 2B (Alibaba, Apache 2.0)
- The performance/RAM champion. 1.9 GB in Q4_K_M, multimodal, 256k context, excellent for short chats and French. Ideal for a simple assistant or voice control.
- Granite 4.2 3B (IBM, Apache 2.0)
- Very lean and token-efficient. 2.2 GB in Q4_K_M. A step up in reasoning quality, built to run H24 without saturating RAM. Strong multilingual performance.
- Qwen 3.5 4B (Alibaba, Apache 2.0)
- The new “small default model” and the most capable of the three. 3.4 GB in Q4_K_M. Fits in RAM on an 8 GB Pi 5, but leaves little headroom—avoid running a heavy GUI alongside it.
#3. Tokens/sec benchmarks (orders of magnitude)
Orders of magnitude measured on a Raspberry Pi 5 8 GB, with the official active cooler, an NVMe SSD via an M.2 HAT, 64-bit Pi OS Bookworm, recent Ollama, Q4_K_M quantization, and a 200-token generation prompt. Figures vary by Ollama version and thermals—treat them as rough estimates, not absolute values.
- Qwen 3.5 2B Q4_K_M
- ≈ 4.4 tokens/sec generation · 6.6 tok/s prompt eval · 1.9 GB RAM · 9.6 W
- Granite 4.2 3B Q4_K_M
- ≈ 3.4 tokens/sec generation · 5.1 tok/s prompt eval · 2.2 GB RAM · 10.2 W
- Qwen 3.5 4B Q4_K_M
- ≈ 2.3 tokens/sec generation · 3.6 tok/s prompt eval · 3.4 GB RAM · 11.2 W
- Gemma 4 E2B (multimodal reference)
- ≈ 3.9 tokens/sec · 4.3 GB RAM · 10.4 W. Fast MoE, but more RAM-hungry.
- Granite 4.2 3B Q3_K_S (efficient variant)
- ≈ 3.6 tokens/sec · 1.7 GB RAM · 9.9 W. The lightest of the bunch, ideal when RAM is limited (4 GB Pi).
The --verbose flag displays the following at the end of generation: eval rate (tok/s), prompt eval rate, and total duration. This is the official metric for comparison.
#4. Relevant offline use cases
At 2-5 tok/s, Pi inference is not intended for real-time chat. It shines when latency is not critical, or when privacy and autonomy matter more than speed.
- Local voice assistant
- Pi 5 + USB microphone + Whisper.cpp tiny + Qwen 3.5 2B + Piper TTS. Question → transcription → response → synthesized voice, all in 8-12 s without the Cloud. Offline home automation, children’s kiosk, educational robot.
- Batch summarization of emails or news
- Cron job that pulls RSS feeds every hour, summarizes them with Qwen 3.5 4B, and sends them to Telegram/Matrix. Inference latency is invisible when it runs asynchronously.
- Log or ticket classification
- Edge Pi that tags incoming logs ('critical error', 'spam', 'normal') before sending them to a central server. Saves WAN bandwidth.
- Embedded / portable demo
- Client presentation, school workshop, conference: Pi 5 in your pocket, USB-C battery, 100% self-contained local AI demo.
- Guardrail for other LLMs
- Small local model that filters/rephrases a prompt before sending it to a cloud LLM. Useful for anonymizing sensitive data at the edge.
#5. Pi-specific optimizations
#Reduce the context window
By default, Ollama uses num_ctx=2048. On Pi, that's already a lot—each context token consumes RAM and slows pre-evaluation. For a short chat assistant, lower it to 1024.
#Limit threads
Ollama uses all cores by default. On Pi 5 (4 cores), this is optimal for raw speed but saturates the SoC, triggering thermal throttling after 1–2 minutes. For long-running use, limiting it to 3 cores maintains a stable throughput for longer.
#Prefer NVMe over microSD
The initial loading of a model from SD reads several GB sequentially. SD A1 → 30–40 s for Phi-3.5 Mini. NVMe → 3–5 s. Once in RAM, inference is identical.
#Drop down one level if you’re short on RAM
On a Pi 5 8 GB with the graphical interface active, Qwen 3.5 4B in Q4_K_M may start swapping. Two options: downgrade to Granite 4.2 3B (2.2 GB) or Qwen 3.5 2B (1.9 GB), freeing 1 to 1.5 GB of RAM without a collapse in quality.
#Troubleshooting
- ollama: command not found after installation
- The systemd script failed (often because Pi OS is too old). Check sudo systemctl status ollama. If you see an 'exec format error', you're on 32-bit—reinstall 64-bit Pi OS.
- The model loads, but at 0.3 tok/s
- You're swapping. Check free -h: if the 'available' column is below 500 MB during inference, your model is too large. Drop down to Granite 4.2 3B or Qwen 3.5 2B, or switch to Q3.
- Thermal throttling after 1 minute
- The SoC exceeded 80 °C. Check vcgencmd measure_temp during inference. Solution: official fan or Argon ONE V3. Without active cooling, the Pi 5 cannot run inference continuously.
- Out of memory halfway through generation
- The Linux OOM killer killed Ollama. Reduce num_ctx to 1024, close the browser, or switch to a smaller model. On a 4 GB Pi, stick with Qwen 3.5 2B, the lightest in the 2026 catalog.
- Service Ollama does not start at boot
- sudo systemctl enable ollama pour activer l'autostart. Vérifiez ensuite avec systemctl is-enabled ollama.
#Go further
Once Ollama is running on your Pi 5, three natural directions to explore further:
- Understand quantization to go further
- The guide to choosing your quantization compares Q3, Q4, Q5, and Q8 in detail. Critical on Pi, where every 100 MB of RAM counts.
- Integrate into a Python app
- The guide on integrating Ollama into Python via the REST API shows how to control your Pi 5 from a Flask/FastAPI script—the foundation of any edge assistant.
- Automate offline workflows
- The guide automate with n8n and Ollama applies as-is on Pi (n8n runs natively on ARM64). Ideal for a standalone AI home automation hub.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.