Vosk: real-time speech recognition, offline ligne
Vosk (alphacep/vosk-api, Apache 2.0 license, developed by Alpha Cephei) is an offline speech-recognition engine that transcribes continuously, with portable models of about 50 MB per language—the lightweight French model vosk-model-small-fr-0.22 weighs 41 MB. It supports 20 or more languages and runs on a Raspberry Pi or smartphone, but remains less accurate than Whisper on difficult audio: reserve it for voice commands and live use.
Vosk is an offline speech recognition engine developed by Alpha Cephei and built on the very lightweight Kaldi engine. It works as a continuous stream: it transcribes as you speak, on a modest processor, using models that are a few dozen megabytes in size. It does not compete with large models on text quality—it does something else, and what it does is not easily replaceable. This guide details the figures published by the project, including those for the two official French models, before making an honest comparison with Whisper.
#What Vosk does differently
Reference transcription models process files: you give them a complete recording, and they return text once processing is finished. Vosk processes a stream: it receives audio in small chunks and produces partial, then final, results as speech unfolds, with near-zero latency thanks to its streaming API—the project's official page describes this “zero-latency response with streaming API” as its central feature, on a par with its model size.
This architectural difference explains everything else: tiny models, effortless CPU execution, a memory footprint compatible with a Raspberry Pi or an entry-level phone, and lower text quality than large models on difficult or noisy audio.
#Vosk in numbers, including French
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
The project is precise about what it offers. The official repository advertises recognition for “20+ languages and dialects,” bindings for Python, Java, Node.js, C#, C++, Rust, and Go, and “portable” 50 MB models, with significantly larger server models for those who want more accuracy. Everything is released under the Apache 2.0 license, a useful guarantee when a project must justify its open-source dependencies to a legal department.
| Model | Size | Word error rate (lower is better) | Intended use |
|---|---|---|---|
| vosk-model-small-fr-0.22 | 41 MB | 23.95 (Common Voice) · 19.30 (MTEDX) · 27.25 (podcasts) | Android, iOS, Raspberry Pi |
| vosk-model-fr-0.22 | 1.4 GB | 14.72 (Common Voice) · 11.64 (MLS) · 13.10 (MTEDX) | Server, best accuracy |
The size difference (41 MB versus 1.4 GB) produces a clear accuracy gap, with the word error rate nearly doubling depending on the test set. This is Vosk’s central trade-off, quantified rather than assumed: the lighter model makes more mistakes; that’s acceptable for a voice command with a limited vocabulary, much less so for transcribing an interview. Beyond the official models, the LinTO project, developed by the French publisher Linagora, publishes its own French models compatible with Vosk—a useful reference for a French-language project that wants to compare multiple sources before choosing.
The details of the four test sets (Common Voice, MTEDX, podcasts, MLS) are also worth reading as a methodological warning rather than as a simple ranking to take at face value: the word error rate ranges from single to double between datasets for the same model, from 19,30 on MTEDX to 27,25 on podcasts for the lightweight version. Clean studio audio and a conversation recorded amid ambient noise do not produce the same quality, regardless of the model chosen—a further reason to test on recordings representative of real-world use instead of relying on a single published average figure on a product page.
#Live browsing, its real specialty
- Latency
- The text appears as you speak, enabling voice commands and live captioning.
- Partial results
- The application can respond before the sentence is finished — essential for a wake word or a short command.
- Footprint
- 41 MB for the lightweight French model, versus 1.4 GB for the server version and often several gigabytes for a large general-purpose model.
- Restricted vocabulary
- You can limit recognition to a list of expected words, which greatly improves accuracy for commands.
This last point is underestimated. To control an application by voice, restricting the vocabulary to fifty possible commands turns an average engine into a highly reliable one—where a large general-purpose model will continue hesitating between similar-sounding homophones. The official specification confirms this as well, listing “reconfigurable vocabulary” among the engine's native features alongside speaker identification—two capabilities that few engines of this size offer natively.
#Vosk or Whisper
| Criterion | Vosk | Whisper and derivatives |
|---|---|---|
| Mode | Continuous flow | File |
| Latency | Virtually none (streaming API) | After processing |
| Model sizes | 41 MB to 1.4 GB depending on the variant | Hundreds of MB to several GB |
| Hardware | Modest, embedded processor | Adequate processor, GPU recommended |
| Quality on difficult audio | Acceptable, measured (WER ≈ 15–28 depending on the model and test) | Significantly better |
| Punctuation and formatting | Limited | Good |
| Choose it for | Voice commands, live, embedded | Transcribing meetings, interviews, and videos |
There is no overall winner. A serious pipeline often uses both: Vosk to detect a command or display immediate feedback while the user speaks, and a heavier model such as Faster-Whisper to produce the final, more accurate transcription once recording is complete and compute is available.
#Cases where it wins
- 01Voice control for an applicationRestricted vocabulary, minimal latency, no network dependencies.
- 02Live captioningInternal meeting, conference, real-time accessibility.
- 03Embedded hardwareA nano-computer or standalone device where a large model won't fit—the project explicitly lists Raspberry Pi and Android as targets.
- 04Strict privacy without a GPUAn ordinary workstation that must not send anything and has no graphics card.
These four use cases share one common trait: they tolerate decent rather than excellent text quality in exchange for near-zero latency and a footprint compatible with modest hardware. As soon as the opposite constraint applies—maximum quality required, latency of a few seconds acceptable—a file-based model such as Faster-Whisper regains the advantage. That is why the two tools most often coexist rather than replace each other in a complete, well-designed voice pipeline.
#Audio formats and text export
In practice, a transcription project is never limited to calling the engine alone: you need to read the received input format, convert it if necessary, then write the result in a format that downstream systems can use, and often handle multiple recording providers at once. The pipeline documented in the study cited above illustrates this complete process: reading several common audio formats (WAV, MP3, FLAC, OGG) through Python preprocessing modules, recognition via Vosk's KaldiRecognizer, then exporting the resulting text to a structured document ready for review or archiving.
This detail matters because Vosk itself expects an audio stream in a specific input format (mono PCM, generally 16 kHz), not just any file as-is. A project receiving heterogeneous audio—phone recordings, videoconferencing exports, compressed files from various sources—therefore almost always needs a conversion step before the engine, and that is often where the recognition-degradation bugs mentioned earlier in this guide hide, rather than in the engine itself, without any explicit error.
#Customize the language model
Under the hood, Vosk recognition relies on its own KaldiRecognizer, derived from the Kaldi speech recognition engine—a detail that explains why the project inherits Kaldi's entire ecosystem of customization tools instead of starting from scratch for every new language or domain. Recent research on the subject, built around a customized Vosk pipeline, measured the concrete effect of an adapted language model: custom models reduce the word error rate, especially for technical vocabulary, strong accents, or noisy audio.
For a French-language project, this opens up a concrete path beyond choosing between the lightweight model and the server model: inject a domain lexicon (product names, internal jargon, acronyms, local proper names) into the language model instead of relying solely on generic vocabulary. It takes more work than a simple download, but it is the lever that comes closest to bringing a Vosk deployment to the quality of a large model in a specific domain, without paying the price in compute or latency.
#Implementation
The process is simple: install the package (pip install vosk in Python), download a model for the desired language, open an audio stream, feed chunks into the engine, and read the partial and then final results. Bindings exist for Python, Java, Node.js, C#, C++, Rust, and Go, and a WebSocket server lets you connect a web page or mobile app to an engine running on your machine.
Two points to keep in mind before going to production: quality depends heavily on the language model you choose—the table above gives you a rough estimate, but test on your own recordings before building on it—and the audio sampling rate must match what the model expects; otherwise, recognition quality degrades without an explicit error message.
- Local voice assistant: the complete pipeline
- Transcribe files quickly on your local machine
- Have the assistant respond aloud with Kokoro
- Whisper + Ollama: 100% offline transcription
- Source: Vosk official GitHub repository
- Source: the Vosk model catalog, with sizes and scores
- Source: study on Vosk language model customization
#FAQ
Is Vosk free?+
Do you need a graphics card?+
Vosk or Whisper?+
Does it work in French, and how well?+
Can you limit the recognized vocabulary?+
What model size should you choose for an embedded project?+
Can you improve accuracy without changing the model?+
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.