Home Assistant + local LLM: private home automation, without cloud
Cloud home automation—Google Home, Alexa, SmartThings—continuously sends snippets of information about what's happening in your home to remote servers. A privacy-focused local home automation LLM reverses the equation: your Home Assistant server controls your lights and thermostats, a Ollama model translates your requests into actions, and nothing leaves the local network. This guide shows how to connect Ollama to Home Assistant in practice, which lightweight model to choose to stay under two seconds of latency, and what you actually gain in privacy.
#Why use a local LLM for a smart home?
With a typical cloud speaker, every sentence you say is sent to a remote server to be transcribed, understood, and acted on. The transcription is retained and sometimes listened to by contractors to train the models, while the manufacturer knows when you get home, what time you turn on the bedroom light, and when you start your evening playlist. For many households, that's too much information given to too many third parties.
Home Assistant addresses the problem at its root: it’s an open-source platform that runs on your hardware and communicates locally with connected devices (Zigbee, Z-Wave, Matter, MQTT). Add a local LLM via Ollama, and you get an assistant that can understand free-form phrases (“lower the living room temperature by two degrees,” “turn off all the downstairs lights”) without any audio or command ever leaving the LAN.
- Real privacy
- No usage data sent to third parties. No advertising profile built from your habits.
- Works offline
- Internet outage, cloud provider failure—the system at home keeps responding to commands.
- Predictable latency
- No more round trips to a datacenter; latency depends only on your machine.
- Full control
- You choose the model, update it whenever you want, and replace it when a better one comes out.
#Prerequisites
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
- Home Assistant in place
- Minimum version 2024.6 (which includes native Ollama integration). Ideally, use the latest stable version. HA OS, HA Container, and HA Supervised all work.
- A machine for Ollama
- Ollama can run on the same machine as HA if it is powerful enough, or on a separate PC on the LAN. A Mac mini M4, an i7 NUC, or a PC with an 8 GB+ GPU can handle it very well.
- Ollama installed
- Daemon active on port 11434. Check with curl http://localhost:11434—the response should be Ollama is running. If nothing is set up, first follow the installation guide Ollama from the site.
- Stable local network
- Ideally, HA and Ollama should be on the same VLAN, with a static IP for the Ollama server. Port 11434 must be open between them.
- A suitable model
- We'll get to that next. For now, remember that you need a model that supports tool calling (calling tools).
#Lightweight models suited to home automation commands
The right model for home automation isn’t necessarily the largest: it must (1) properly support tool calling to call Home Assistant services, (2) speak French well, and (3) respond quickly. 3B to 8B models in Q4_K_M are the sweet spot.
- Qwen 3.5 9B (Q4_K_M, ~6.6 GB VRAM)
- The reference 8 GB choice in 2026. Robust tool calling, excellent French, vision, and 256k context, with reasonable GPU latency. It's the right starting point for home automation.
- Granite 4.2 8B (Q4_K_M, ~5.3 GB VRAM)
- IBM, Apache 2.0, very token-efficient with 128k context. Reliable, restrained tool calling: a good alternative if Qwen drifts on your commands.
- Granite 4.2 3B (Q4_K_M, ~2.2 GB VRAM)
- For modest setups (Pi 5, NUC without a GPU). Very lightweight, with tool calling available but less reliable for complex commands—stick to short, explicit sentences.
- Gemma 4 12B (Q4_K_M, ~7.6 GB VRAM)
- If you have a RTX 3060 12 GB or better. Multimodal, Apache 2.0, with a clear improvement on ambiguous commands (“set a cozy mood in the living room”) and contextual understanding.
- Mistral Small 24B (Q4_K_M, ~14 GB VRAM)
- Very good in French and strong as a general-purpose model; best reserved for 16 GB cards and larger. Comfortable for rich dialogue in addition to strict home control.
#1. Connect Ollama to Home Assistant
The Ollama integration has been officially supported in Home Assistant since version 2024.6. It can be configured entirely through the interface, without touching configuration.yaml.
- 01Move the model to OllamaOn the machine hosting Ollama, run the shell command below. The download takes 3-5 GB depending on the selected model.
- 01Make Ollama accessible on the LANBy default, Ollama listens only on localhost. To let Home Assistant (potentially running on another machine) call it, expose the service on all interfaces.
- 01Add the integration to Home AssistantSettings > Devices & services > Add integration > search for “Ollama”.
- 02Enter the URLURL: http://IP-DU-SERVEUR-OLLAMA:11434 (e.g., http://192.168.1.42:11434). If Ollama runs on the same machine as HA Container, use http://host.docker.internal:11434.
- 03Choose the modelHome Assistant automatically queries Ollama and displays the list of available models. Select qwen3.5:9b.
- 04Enable control (Home Assistant Control)In the integration options, check “Control Home Assistant.” This authorizes the LLM to call services (turn on, turn off, adjust the temperature). Without this option, it can only chat.
#2. Expose the right entities to the assistant
Home Assistant sometimes manages hundreds of entities (sensors, plugs, bulbs, views, automations…). Sending them all to the LLM would saturate the context and confuse the model. The trick is to expose only what the assistant actually needs to control.
- 01Settings > Voice > ExposeList of all entities. Enable the ones the assistant should control: lights, thermostats, outlets, shutters. Disable internal sensors (Pi RAM, Zigbee signal, and so on) that are useless to a human.
- 02Rename in natural languageAn entity named light.salon_lampe_principale_zigbee_3 is unreadable. Rename it to “Lampe principale du salon.” The LLM chooses the right entity much more reliably from a human-readable name.
- 03Group by zoneAssign each entity to a zone (Living Room, Kitchen, Bedroom…). The LLM can then understand “turn off the living room” without manually listing every lamp.
- 04TestIn Settings > Voice > Assistants, click the chat icon for your Ollama pipeline and try: “What is the room temperature?” followed by “Turn on the living room’s main light.”
#3. 100% local voice pipeline (optional)
If you want to speak aloud (instead of typing in the HA app), Assist can chain speech recognition → LLM → speech synthesis, entirely locally. No Google or Amazon dependency.
- Wake word — openWakeWord
- Official HA add-on. Detects an activation keyword (“Hey Jarvis”, “Nabu”) on a Pi or a satellite ESP32 (Atom Echo, M5Stack).
- STT (speech-to-text) — Whisper
- Official HA add-on based on faster-whisper. The tiny or base model is enough for short commands in French. ~1-2 s transcription on a decent CPU.
- LLM — Ollama (already configured)
- Receives the transcript, calls the appropriate HA tool, and returns a confirmation sentence.
- TTS (text-to-speech) — Piper
- Official HA add-on. Native French voices (fr_FR-siwis-medium, fr_FR-tom-medium). Very fast, ~200 ms per sentence.
Once the four add-ons are installed and started, Settings > Voice > Assistants lets you assemble the pipeline: select Whisper for STT, your Ollama integration as the conversational agent, and Piper for TTS. This is when your home becomes voice-controllable, with no outbound connection.
#Latency: local vs. cloud
This is the often-overlooked argument. The cloud is not magic: every request makes a round trip to the data center, plus a queue on the server side. Locally, you cut the journey in half.
- Cloud (Google/Alexa)
- Cloud STT (~400 ms) + cloud intent (~300 ms) + HA action via cloud (~200 ms) + TTS (~300 ms) = 1.2–1.8 s on average, longer when the network is busy.
- Local — Pi 5 + 3B CPU LLM
- STT Whisper tiny (~800 ms) + 3B LLM (~3-5 s on CPU) + action (~50 ms) + Piper TTS (~200 ms) = 4-6 s. Acceptable for a command, frustrating for intensive use.
- Local — PC + 8B LLM GPU 8 GB
- Whisper base (~400 ms) + 8B Q4 LLM (~600–900 ms) + action (~50 ms) + TTS (~200 ms) = 1.3–1.6 s. As fast as a Google Home, without the cloud.
- Local — Mac mini M4 24 GB
- Full pipeline ~1.1–1.4 s thanks to the integrated GPU and unified memory. This is probably the best performance-per-power ratio for 24/7 LLM home automation.
#What really leaves the LAN (nothing)
Let’s concretely check the scope of the data. With a properly configured HA + Ollama + Whisper + Piper setup:
- Audio captured by the microphone
- Processed by local Whisper. No audio leaves the network.
- Transcribed text from your command
- Sent to local Ollama. No transcription leaves the network.
- List of your devices and zones
- Shared with Ollama locally so it can call the right entity. Stays on the LAN.
- HA decision and action
- Executed by HA, which communicates directly with Zigbee/Z-Wave/Matter/MQTT locally.
- Piper voice response
- Audio synthesized locally. No synthetic voice fetched online.
The only possible outbound traffic is still HA, Ollama, and model updates—you decide when to run them. For strict use, you can cut off Internet access to the HA machine after installation: everything continues to work.
#Tips and troubleshooting
- The LLM responds in English
- Add a system prompt to the Ollama integration: “You are the home's home-automation assistant. Always reply in French, in one short sentence.” This instruction is sufficient with Qwen 3.5 or Granite 4.2.
- It invents nonexistent entities
- A symptom of a model that is too small or an ambiguous entity name. Rename the entities in natural language, reduce the number exposed, or switch to Gemma 4 12B if VRAM is available.
- It calls the right service but in the wrong region
- Make sure every entity has an “area” assigned in HA. Without a zone, the model guesses and gets similar names wrong.
- Latency collapses after a few hours
- Ollama unloaded the model. Start the service with OLLAMA_KEEP_ALIVE=24h (or 168h for a week).
- Connection refused from HA
- Ollama is still listening on 127.0.0.1. Check OLLAMA_HOST=0.0.0.0:11434 in the systemd unit, restart, and test with curl from the HA machine.
- Pi 5 overheating under CPU load
- The CPU LLM pushes all 4 cores to 100%. Add an active fan (or an Argon ONE V3 Pi 5 case) if you want to run everything on the Pi. Otherwise, offload Ollama to another host.
- Want to confine Ollama to the LAN only
- Rather than 0.0.0.0:11434, listen on the specific LAN IP (e.g., OLLAMA_HOST=192.168.1.42:11434). Ollama will no longer accept connections from other interfaces.
#Go further
You have a working local LLM home automation assistant. Here are a few ways to take it further:
- Install Ollama on Linux
- If you want to dedicate a Linux machine to Ollama (the cleanest setup for 24/7 home automation), the Linux installation guide covers systemd, NVIDIA/AMD GPUs, and network configuration.
- LLM on Raspberry Pi 5
- To assess concretely what runs on pure CPU on a Pi 5 8 GB, with tokens/sec benchmarks and a selection of 1-3B models.
- Choosing your quantization (Q4, Q5, Q8)
- To understand why Q4_K_M is the smart-home sweet spot, and when moving up to Q5_K_M or down to Q3 makes sense.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.