100% local voice assistant: Whisper + Ollama + Piper
A local voice assistant listens, understands, and responds aloud without ever sending a byte to the cloud. This guide combines three open-source components—Whisper for speech recognition, an LLM served by Ollama for reasoning, and Piper for a natural French voice—into a complete loop running on your own machine. We explain the STT → LLM → TTS pipeline, recommended hardware, and the latency you can realistically expect.
#Why use a local voice assistant
Alexa, Google Assistant, and Siri share the same flaw: every sentence you speak is sent to remote servers to be transcribed and interpreted. A local voice assistant reverses this model. The microphone, recognition, reasoning, and response voice all run on your hardware. Nothing travels over the internet, and the assistant keeps working even when the connection is down.
Beyond privacy, the benefit is control. You choose the language model, its personality through the system prompt, and the output voice, and you can connect it to your own tools—home automation, calendar, notes—without depending on a third-party API that could shut down or change its terms overnight.
- Privacy
- No voice command is recorded or analyzed by a third party. Ideal for sensitive home or professional use.
- Hors-ligne
- The complete pipeline works without internet access once the models have been downloaded.
- Customizable
- System prompt, voice, language, wake word, and tools are all under your control.
- No subscription
- Everything is open source. Your only costs are your hardware and electricity.
#The STT → LLM → TTS pipeline
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
A voice assistant, no matter how sophisticated, is just a sequence of three steps. Understanding this breakdown is the key to diagnosing latency and replacing one link without breaking the rest.
- 01STT — Speech To TextThe microphone captures your voice, and Whisper turns it into text. This is speech recognition. Quality depends on the Whisper model you choose and the ambient noise.
- 02LLM — reasoningThe transcribed text is sent to a language model served by Ollama, with a system prompt that defines the assistant's role. The LLM generates a written response.
- 03TTS — Text To SpeechThe text response was passed to Piper, which reads it aloud with a French voice. This is text-to-speech.
Perceived latency is the sum of three stages: Whisper transcription time, LLM generation time (the most variable), and Piper synthesis time. We’ll revisit this in detail below.
#Requirements and hardware
A local voice assistant is more demanding than a simple text chat: Whisper and the LLM share the GPU, and responsiveness matters. Here are the hardware guidelines based on your goals.
- Minimal (CPU only)
- Whisper base plus a small LLM in Q4_K_M (≈2 GB), such as Qwen 3.5 2B or Granite 4.2 3B. It works, but expect several seconds of latency per turn. Acceptable for home automation where you do not speak continuously.
- Comfortable
- RTX 3060 12GB or RTX 4070 12GB. Whisper small/medium on the GPU plus an 8–9B Q4_K_M LLM (≈5–7 GB), such as Qwen 3.5 9B or Granite 4.2 8B. Latency is about 1 to 3 seconds per turn.
- Smooth
- RTX 4080 16GB / RTX 4090 24GB or a Mac M4 Pro (24–48 GB unified). Whisper medium + a more capable LLM (Gemma 4 12B ≈8 GB, or Qwen 3.8 27B ≈18 GB on 24 GB), with near-instant responses.
- Microphone and audio
- A decent USB microphone is sufficient. On Linux, use PipeWire or PulseAudio; on macOS/Windows, the default configuration works.
On the software side, we assume Ollama is already installed and working. If not, start with the Ollama installation guide before continuing. Verify that the daemon responds.
#1. Listening and transcription (Whisper)
Whisper is OpenAI's open-source speech recognition model. For real-time local use, faster-whisper is preferred: an optimized reimplementation that uses CTranslate2 and runs much faster than the reference version, on both CPU and GPU.
Whisper models come in several sizes. The larger the model, the better the transcription—especially in French and noisy environments—but the slower it is.
- tiny / base
- Very fast, perfect for a modest CPU or home automation. Accurate transcription of short commands in French.
- small
- A good speed/quality compromise for most assistants. Recommended as a starting point.
- medium
- Significantly better with long sentences and accents, but requires a GPU to remain responsive.
- large-v3
- Maximum quality, reserved for capable GPUs when latency matters.
Here is a listening function that records from the microphone and returns the transcribed text. We force French as the language to avoid detection errors.
#2. Reasoning (Ollama)
The transcribed text now goes to the LLM. Ollama exposes an HTTP API at http://localhost:11434; we query it in Python. The model choice depends on your GPU: a compact model such as Qwen 3.5 9B (6.6 GB, 256k context, Apache 2.0) or Granite 4.2 8B (5.3 GB) in Q4_K_M offers an excellent balance for a conversational assistant.
The system prompt is what turns a generic LLM into a useful voice assistant. For voice, one crucial instruction is to request short answers. A wall of text becomes endless when read aloud.
#3. French speech synthesis (Piper)
Piper is a fast, local neural text-to-speech engine designed by the Home Assistant community (Rhasspy project). It produces a much more natural French voice than older engines such as espeak, while remaining lightweight enough to run on a Raspberry Pi.
Piper uses pretrained voices, one per language and voice type. For French, several voices are available (fr_FR) in different qualities. Each voice is a pair of files: the .onnx model and its .onnx.json configuration.
It is then used to turn a sentence into audio and play it directly.
#4. Assemble the complete loop
With the three building blocks tested separately, all that's left is to chain them in a loop: listen, reason, speak, repeat. Here is the complete minimal assistant, bringing together the preceding functions.
That's it: about a hundred lines spread across three files, and you have a voice assistant that listens in French, reasons with a local LLM, and responds aloud, without a single outgoing network call. From there, you can add a wake word (“Ok maison”) with openWakeWord, a VAD for continuous listening, or connect tools.
#Real-world latency and optimizations
The question that makes or breaks a voice assistant: how long between the end of your sentence and the start of the spoken response? Here are approximate figures observed on a RTX 4070 with Whisper small and an 8–9B Q4_K_M model (such as Qwen 3.5 9B).
- Whisper transcription
- ≈ 0.3 to 0.8 s for a 5-second sentence with the small model on GPU. On CPU, expect 2 to 4 s.
- LLM generation
- The most variable position. An 8-9B model on a GPU produces its first response in ≈ 0.5 to 1.5 s depending on the length. This is where a low num_predict value and a fast model matter most.
- Piper summary
- ≈ 0.2 to 0.5 s for one or two sentences with a medium voice. Very fast, rarely the bottleneck.
- Perceived total
- Around 1 to 3 seconds per turn on a mid-range GPU. Under one second requires high-end hardware or streaming.
- Keep models loaded
- Set OLLAMA_KEEP_ALIVE to prevent Ollama from unloading the model between questions — reloading takes several seconds.
- Short responses
- A low num_predict value and a system prompt that requires concision directly reduce generation time and reading time.
- Adapted Whisper
- Don’t choose large-v3 if small is sufficient for your audio: you save hundreds of milliseconds on every turn.
#5. Integrate with Home Assistant
If your goal is to control a smart home by voice, you do not need to recode everything. Home Assistant natively integrates the "Assist" pipeline, and Whisper and Piper can be connected through official add-ons—they were in fact developed within this ecosystem.
- 01Install voice add-onsFrom Settings → Add-ons, add “Whisper” and “Piper.” Home Assistant runs them locally as STT and TTS services.
- 02Create an Assist pipelineIn Settings → Voice, create an assistant: choose Whisper for recognition and Piper (fr_FR voice) for synthesis.
- 03Connect the Ollama LLMHome Assistant offers a “Ollama” integration as a conversational agent. Enter the http://localhost:11434 URL and your model to replace the default intent analysis.
- 04Choose a voice entry pointAn ESP32 satellite (Voice PE), an old phone, or the mobile app serves as a microphone/speaker in the rooms.
#Troubleshooting
- Whisper transcribes in the wrong language
- Force language='fr' in transcribe(). Without it, a short sentence may be detected as English.
- No audio output
- Check sounddevice's output device (sd.query_devices()) and make sure the samplerate passed to sd.play matches voix.config.sample_rate.
- The microphone records nothing
- Run sd.query_devices() to find the correct input index, and check microphone permissions (macOS) or the PipeWire/PulseAudio source (Linux).
- Responses that are too long and unreadable
- Strengthen the system prompt's brevity instruction and lower num_predict. A verbose LLM ruins the voice experience.
- Latency increasing throughout the conversation
- History bloats the context. Truncate it to the last N messages.
- Ollama reloads the model for every question
- Increase OLLAMA_KEEP_ALIVE (for example, to 30m) to keep the model in VRAM between turns.
#Go further
Your voice assistant relies on a well-chosen, well-tuned local LLM. These site guides complete the setup:
- Whisper + Ollama: transcription and summarization
- To explore the audio side in greater depth and process long files rather than real time.
- Local LLM for home automation
- To connect your assistant to Home Assistant and control your home by voice while preserving privacy.
- Q4, Q5, Q8: which quantization should you choose
- To balance response quality, speed, and VRAM based on your GPU.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.