Intermediate 11 minHome automation

Home Assistant + local LLM: private home automation, without cloud

Cloud home automation—Google Home, Alexa, SmartThings—continuously sends snippets of information about what's happening in your home to remote servers. A privacy-focused local home automation LLM reverses the equation: your Home Assistant server controls your lights and thermostats, a Ollama model translates your requests into actions, and nothing leaves the local network. This guide shows how to connect Ollama to Home Assistant in practice, which lightweight model to choose to stay under two seconds of latency, and what you actually gain in privacy.

By Mohamed Meguedmi·Update 2026-08-27·Tested on Windows, macOS, and Linux

#Why use a local LLM for a smart home?

With a typical cloud speaker, every sentence you say is sent to a remote server to be transcribed, understood, and acted on. The transcription is retained and sometimes listened to by contractors to train the models, while the manufacturer knows when you get home, what time you turn on the bedroom light, and when you start your evening playlist. For many households, that's too much information given to too many third parties.

Home Assistant addresses the problem at its root: it’s an open-source platform that runs on your hardware and communicates locally with connected devices (Zigbee, Z-Wave, Matter, MQTT). Add a local LLM via Ollama, and you get an assistant that can understand free-form phrases (“lower the living room temperature by two degrees,” “turn off all the downstairs lights”) without any audio or command ever leaving the LAN.

Real privacy
No usage data sent to third parties. No advertising profile built from your habits.
Works offline
Internet outage, cloud provider failure—the system at home keeps responding to commands.
Predictable latency
No more round trips to a datacenter; latency depends only on your machine.
Full control
You choose the model, update it whenever you want, and replace it when a better one comes out.
i
Home Assistant in two sentences
This is a Python server that brings together hundreds of home automation integrations under a single interface. It deploys in minutes on a Raspberry Pi, NUC, mini-PC, or Docker container—and does not depend on any online service to operate.

#Prerequisites

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates
Home Assistant in place
Minimum version 2024.6 (which includes native Ollama integration). Ideally, use the latest stable version. HA OS, HA Container, and HA Supervised all work.
A machine for Ollama
Ollama can run on the same machine as HA if it is powerful enough, or on a separate PC on the LAN. A Mac mini M4, an i7 NUC, or a PC with an 8 GB+ GPU can handle it very well.
Ollama installed
Daemon active on port 11434. Check with curl http://localhost:11434—the response should be Ollama is running. If nothing is set up, first follow the installation guide Ollama from the site.
Stable local network
Ideally, HA and Ollama should be on the same VLAN, with a static IP for the Ollama server. Port 11434 must be open between them.
A suitable model
We'll get to that next. For now, remember that you need a model that supports tool calling (calling tools).
!
Is the Raspberry Pi 5 alone enough?
A Pi 5 8 GB can run Home Assistant AND a small 3B model on the CPU for short commands—but latency rises quickly (4–8 s per request). For a smooth family setup, dedicate one machine to Ollama and keep the Pi for HA. See the LLM on Raspberry Pi 5 guide if you want to combine everything.

#Lightweight models suited to home automation commands

The right model for home automation isn’t necessarily the largest: it must (1) properly support tool calling to call Home Assistant services, (2) speak French well, and (3) respond quickly. 3B to 8B models in Q4_K_M are the sweet spot.

Qwen 3.5 9B (Q4_K_M, ~6.6 GB VRAM)
The reference 8 GB choice in 2026. Robust tool calling, excellent French, vision, and 256k context, with reasonable GPU latency. It's the right starting point for home automation.
Granite 4.2 8B (Q4_K_M, ~5.3 GB VRAM)
IBM, Apache 2.0, very token-efficient with 128k context. Reliable, restrained tool calling: a good alternative if Qwen drifts on your commands.
Granite 4.2 3B (Q4_K_M, ~2.2 GB VRAM)
For modest setups (Pi 5, NUC without a GPU). Very lightweight, with tool calling available but less reliable for complex commands—stick to short, explicit sentences.
Gemma 4 12B (Q4_K_M, ~7.6 GB VRAM)
If you have a RTX 3060 12 GB or better. Multimodal, Apache 2.0, with a clear improvement on ambiguous commands (“set a cozy mood in the living room”) and contextual understanding.
Mistral Small 24B (Q4_K_M, ~14 GB VRAM)
Very good in French and strong as a general-purpose model; best reserved for 16 GB cards and larger. Comfortable for rich dialogue in addition to strict home control.
→
Recommended quantization: Q4_K_M
For home automation, Q4_K_M offers the best VRAM/quality tradeoff. There is no need to move up to Q5 or Q8: for short, structured commands, the difference is invisible. Reserve higher quantizations for open-ended dialogue or RAG.

#1. Connect Ollama to Home Assistant

The Ollama integration has been officially supported in Home Assistant since version 2024.6. It can be configured entirely through the interface, without touching configuration.yaml.

  1. 01
    Move the model to Ollama
    On the machine hosting Ollama, run the shell command below. The download takes 3-5 GB depending on the selected model.
Terminal — machine Ollama
ollama pull qwen3.5:9b
ollama list  # vérifier que le modèle apparaît
  1. 01
    Make Ollama accessible on the LAN
    By default, Ollama listens only on localhost. To let Home Assistant (potentially running on another machine) call it, expose the service on all interfaces.
Linux/macOS — expose Ollama on the LAN
# systemd (Linux)
sudo systemctl edit ollama
# ajouter dans le bloc [Service]
# Environment="OLLAMA_HOST=0.0.0.0:11434"
sudo systemctl restart ollama

# macOS (lancement manuel ou launchctl)
launchctl setenv OLLAMA_HOST "0.0.0.0:11434"
# puis relancer Ollama
  1. 01
    Add the integration to Home Assistant
    Settings > Devices & services > Add integration > search for “Ollama”.
  2. 02
    Enter the URL
    URL: http://IP-DU-SERVEUR-OLLAMA:11434 (e.g., http://192.168.1.42:11434). If Ollama runs on the same machine as HA Container, use http://host.docker.internal:11434.
  3. 03
    Choose the model
    Home Assistant automatically queries Ollama and displays the list of available models. Select qwen3.5:9b.
  4. 04
    Enable control (Home Assistant Control)
    In the integration options, check “Control Home Assistant.” This authorizes the LLM to call services (turn on, turn off, adjust the temperature). Without this option, it can only chat.
!
No authentication by default on Ollama
Ollama has no built-in authentication. If you expose port 11434 on the LAN, make sure your router does not route it to the Internet. For a tighter setup, put Ollama behind a reverse proxy with basic auth (Caddy, nginx), or restrict it by IP at the firewall.

#2. Expose the right entities to the assistant

Home Assistant sometimes manages hundreds of entities (sensors, plugs, bulbs, views, automations…). Sending them all to the LLM would saturate the context and confuse the model. The trick is to expose only what the assistant actually needs to control.

  1. 01
    Settings > Voice > Expose
    List of all entities. Enable the ones the assistant should control: lights, thermostats, outlets, shutters. Disable internal sensors (Pi RAM, Zigbee signal, and so on) that are useless to a human.
  2. 02
    Rename in natural language
    An entity named light.salon_lampe_principale_zigbee_3 is unreadable. Rename it to “Lampe principale du salon.” The LLM chooses the right entity much more reliably from a human-readable name.
  3. 03
    Group by zone
    Assign each entity to a zone (Living Room, Kitchen, Bedroom…). The LLM can then understand “turn off the living room” without manually listing every lamp.
  4. 04
    Test
    In Settings > Voice > Assistants, click the chat icon for your Ollama pipeline and try: “What is the room temperature?” followed by “Turn on the living room’s main light.”
→
Reasonable limit: 30–50 exposed entities
Beyond that, 7–8B models start confusing nearby entities (“bedside lamp” vs. “desk lamp”). If you have 200 light bulbs, use HA areas and groups instead of exposing each entity individually.

#3. 100% local voice pipeline (optional)

If you want to speak aloud (instead of typing in the HA app), Assist can chain speech recognition → LLM → speech synthesis, entirely locally. No Google or Amazon dependency.

Wake word — openWakeWord
Official HA add-on. Detects an activation keyword (“Hey Jarvis”, “Nabu”) on a Pi or a satellite ESP32 (Atom Echo, M5Stack).
STT (speech-to-text) — Whisper
Official HA add-on based on faster-whisper. The tiny or base model is enough for short commands in French. ~1-2 s transcription on a decent CPU.
LLM — Ollama (already configured)
Receives the transcript, calls the appropriate HA tool, and returns a confirmation sentence.
TTS (text-to-speech) — Piper
Official HA add-on. Native French voices (fr_FR-siwis-medium, fr_FR-tom-medium). Very fast, ~200 ms per sentence.

Once the four add-ons are installed and started, Settings > Voice > Assistants lets you assemble the pipeline: select Whisper for STT, your Ollama integration as the conversational agent, and Piper for TTS. This is when your home becomes voice-controllable, with no outbound connection.

i
Voice satellite without a PC microphone
An ESP32-S3-BOX or an M5Stack Atom Echo flashed with ESPHome becomes an Assist satellite: it detects the wake word, sends audio to HA, and plays the TTS response. Cost: €20–30. Multiple satellites can coexist in the same home.

#Latency: local vs. cloud

This is the often-overlooked argument. The cloud is not magic: every request makes a round trip to the data center, plus a queue on the server side. Locally, you cut the journey in half.

Cloud (Google/Alexa)
Cloud STT (~400 ms) + cloud intent (~300 ms) + HA action via cloud (~200 ms) + TTS (~300 ms) = 1.2–1.8 s on average, longer when the network is busy.
Local — Pi 5 + 3B CPU LLM
STT Whisper tiny (~800 ms) + 3B LLM (~3-5 s on CPU) + action (~50 ms) + Piper TTS (~200 ms) = 4-6 s. Acceptable for a command, frustrating for intensive use.
Local — PC + 8B LLM GPU 8 GB
Whisper base (~400 ms) + 8B Q4 LLM (~600–900 ms) + action (~50 ms) + TTS (~200 ms) = 1.3–1.6 s. As fast as a Google Home, without the cloud.
Local — Mac mini M4 24 GB
Full pipeline ~1.1–1.4 s thanks to the integrated GPU and unified memory. This is probably the best performance-per-power ratio for 24/7 LLM home automation.
→
Keep the model warm
Ollama unloads inactive models after 5 minutes by default. For home automation, start Ollama with OLLAMA_KEEP_ALIVE=24h—the model stays in VRAM, and the first command after several hours of silence responds as quickly as the hundredth.

#What really leaves the LAN (nothing)

Let’s concretely check the scope of the data. With a properly configured HA + Ollama + Whisper + Piper setup:

Audio captured by the microphone
Processed by local Whisper. No audio leaves the network.
Transcribed text from your command
Sent to local Ollama. No transcription leaves the network.
List of your devices and zones
Shared with Ollama locally so it can call the right entity. Stays on the LAN.
HA decision and action
Executed by HA, which communicates directly with Zigbee/Z-Wave/Matter/MQTT locally.
Piper voice response
Audio synthesized locally. No synthetic voice fetched online.

The only possible outbound traffic is still HA, Ollama, and model updates—you decide when to run them. For strict use, you can cut off Internet access to the HA machine after installation: everything continues to work.

i
To verify it yourself
On your router (or with tcpdump on the host), filter outgoing traffic from the HA and Ollama IPs while using the assistant. You will see only DNS resolutions and possibly NTP—no traffic to any third-party AI service.

#Tips and troubleshooting

The LLM responds in English
Add a system prompt to the Ollama integration: “You are the home's home-automation assistant. Always reply in French, in one short sentence.” This instruction is sufficient with Qwen 3.5 or Granite 4.2.
It invents nonexistent entities
A symptom of a model that is too small or an ambiguous entity name. Rename the entities in natural language, reduce the number exposed, or switch to Gemma 4 12B if VRAM is available.
It calls the right service but in the wrong region
Make sure every entity has an “area” assigned in HA. Without a zone, the model guesses and gets similar names wrong.
Latency collapses after a few hours
Ollama unloaded the model. Start the service with OLLAMA_KEEP_ALIVE=24h (or 168h for a week).
Connection refused from HA
Ollama is still listening on 127.0.0.1. Check OLLAMA_HOST=0.0.0.0:11434 in the systemd unit, restart, and test with curl from the HA machine.
Pi 5 overheating under CPU load
The CPU LLM pushes all 4 cores to 100%. Add an active fan (or an Argon ONE V3 Pi 5 case) if you want to run everything on the Pi. Otherwise, offload Ollama to another host.
Want to confine Ollama to the LAN only
Rather than 0.0.0.0:11434, listen on the specific LAN IP (e.g., OLLAMA_HOST=192.168.1.42:11434). Ollama will no longer accept connections from other interfaces.

#Go further

You have a working local LLM home automation assistant. Here are a few ways to take it further:

Install Ollama on Linux
If you want to dedicate a Linux machine to Ollama (the cleanest setup for 24/7 home automation), the Linux installation guide covers systemd, NVIDIA/AMD GPUs, and network configuration.
LLM on Raspberry Pi 5
To assess concretely what runs on pure CPU on a Pi 5 8 GB, with tokens/sec benchmarks and a selection of 1-3B models.
Choosing your quantization (Q4, Q5, Q8)
To understand why Q4_K_M is the smart-home sweet spot, and when moving up to Q5_K_M or down to Q3 makes sense.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.

Prices in euros (€) are French market prices including VAT, as checked by BestLLMfor. US prices differ: the Amazon buttons show the current US price.