Intermediate 11 minAudio

Vosk: real-time speech recognition, offline ligne

Direct response

Vosk (alphacep/vosk-api, Apache 2.0 license, developed by Alpha Cephei) is an offline speech-recognition engine that transcribes continuously, with portable models of about 50 MB per language—the lightweight French model vosk-model-small-fr-0.22 weighs 41 MB. It supports 20 or more languages and runs on a Raspberry Pi or smartphone, but remains less accurate than Whisper on difficult audio: reserve it for voice commands and live use.

Vosk is an offline speech recognition engine developed by Alpha Cephei and built on the very lightweight Kaldi engine. It works as a continuous stream: it transcribes as you speak, on a modest processor, using models that are a few dozen megabytes in size. It does not compete with large models on text quality—it does something else, and what it does is not easily replaceable. This guide details the figures published by the project, including those for the two official French models, before making an honest comparison with Whisper.

By Mohamed Meguedmi·Update 2026-09-28·Tested on Windows, macOS, and Linux

#What Vosk does differently

Reference transcription models process files: you give them a complete recording, and they return text once processing is finished. Vosk processes a stream: it receives audio in small chunks and produces partial, then final, results as speech unfolds, with near-zero latency thanks to its streaming API—the project's official page describes this “zero-latency response with streaming API” as its central feature, on a par with its model size.

This architectural difference explains everything else: tiny models, effortless CPU execution, a memory footprint compatible with a Raspberry Pi or an entry-level phone, and lower text quality than large models on difficult or noisy audio.

#Vosk in numbers, including French

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

The project is precise about what it offers. The official repository advertises recognition for “20+ languages and dialects,” bindings for Python, Java, Node.js, C#, C++, Rust, and Go, and “portable” 50 MB models, with significantly larger server models for those who want more accuracy. Everything is released under the Apache 2.0 license, a useful guarantee when a project must justify its open-source dependencies to a legal department.

The two official French models (alphacephei.com/vosk/models catalog)
ModelSizeWord error rate (lower is better)Intended use
vosk-model-small-fr-0.2241 MB23.95 (Common Voice) · 19.30 (MTEDX) · 27.25 (podcasts)Android, iOS, Raspberry Pi
vosk-model-fr-0.221.4 GB14.72 (Common Voice) · 11.64 (MLS) · 13.10 (MTEDX)Server, best accuracy

The size difference (41 MB versus 1.4 GB) produces a clear accuracy gap, with the word error rate nearly doubling depending on the test set. This is Vosk’s central trade-off, quantified rather than assumed: the lighter model makes more mistakes; that’s acceptable for a voice command with a limited vocabulary, much less so for transcribing an interview. Beyond the official models, the LinTO project, developed by the French publisher Linagora, publishes its own French models compatible with Vosk—a useful reference for a French-language project that wants to compare multiple sources before choosing.

The details of the four test sets (Common Voice, MTEDX, podcasts, MLS) are also worth reading as a methodological warning rather than as a simple ranking to take at face value: the word error rate ranges from single to double between datasets for the same model, from 19,30 on MTEDX to 27,25 on podcasts for the lightweight version. Clean studio audio and a conversation recorded amid ambient noise do not produce the same quality, regardless of the model chosen—a further reason to test on recordings representative of real-world use instead of relying on a single published average figure on a product page.

#Live browsing, its real specialty

Latency
The text appears as you speak, enabling voice commands and live captioning.
Partial results
The application can respond before the sentence is finished — essential for a wake word or a short command.
Footprint
41 MB for the lightweight French model, versus 1.4 GB for the server version and often several gigabytes for a large general-purpose model.
Restricted vocabulary
You can limit recognition to a list of expected words, which greatly improves accuracy for commands.

This last point is underestimated. To control an application by voice, restricting the vocabulary to fifty possible commands turns an average engine into a highly reliable one—where a large general-purpose model will continue hesitating between similar-sounding homophones. The official specification confirms this as well, listing “reconfigurable vocabulary” among the engine's native features alongside speaker identification—two capabilities that few engines of this size offer natively.

#Vosk or Whisper

Two tools, two jobs
CriterionVoskWhisper and derivatives
ModeContinuous flowFile
LatencyVirtually none (streaming API)After processing
Model sizes41 MB to 1.4 GB depending on the variantHundreds of MB to several GB
HardwareModest, embedded processorAdequate processor, GPU recommended
Quality on difficult audioAcceptable, measured (WER ≈ 15–28 depending on the model and test)Significantly better
Punctuation and formattingLimitedGood
Choose it forVoice commands, live, embeddedTranscribing meetings, interviews, and videos

There is no overall winner. A serious pipeline often uses both: Vosk to detect a command or display immediate feedback while the user speaks, and a heavier model such as Faster-Whisper to produce the final, more accurate transcription once recording is complete and compute is available.

#Cases where it wins

  1. 01
    Voice control for an application
    Restricted vocabulary, minimal latency, no network dependencies.
  2. 02
    Live captioning
    Internal meeting, conference, real-time accessibility.
  3. 03
    Embedded hardware
    A nano-computer or standalone device where a large model won't fit—the project explicitly lists Raspberry Pi and Android as targets.
  4. 04
    Strict privacy without a GPU
    An ordinary workstation that must not send anything and has no graphics card.

These four use cases share one common trait: they tolerate decent rather than excellent text quality in exchange for near-zero latency and a footprint compatible with modest hardware. As soon as the opposite constraint applies—maximum quality required, latency of a few seconds acceptable—a file-based model such as Faster-Whisper regains the advantage. That is why the two tools most often coexist rather than replace each other in a complete, well-designed voice pipeline.

#Audio formats and text export

In practice, a transcription project is never limited to calling the engine alone: you need to read the received input format, convert it if necessary, then write the result in a format that downstream systems can use, and often handle multiple recording providers at once. The pipeline documented in the study cited above illustrates this complete process: reading several common audio formats (WAV, MP3, FLAC, OGG) through Python preprocessing modules, recognition via Vosk's KaldiRecognizer, then exporting the resulting text to a structured document ready for review or archiving.

This detail matters because Vosk itself expects an audio stream in a specific input format (mono PCM, generally 16 kHz), not just any file as-is. A project receiving heterogeneous audio—phone recordings, videoconferencing exports, compressed files from various sources—therefore almost always needs a conversion step before the engine, and that is often where the recognition-degradation bugs mentioned earlier in this guide hide, rather than in the engine itself, without any explicit error.

#Customize the language model

Under the hood, Vosk recognition relies on its own KaldiRecognizer, derived from the Kaldi speech recognition engine—a detail that explains why the project inherits Kaldi's entire ecosystem of customization tools instead of starting from scratch for every new language or domain. Recent research on the subject, built around a customized Vosk pipeline, measured the concrete effect of an adapted language model: custom models reduce the word error rate, especially for technical vocabulary, strong accents, or noisy audio.

For a French-language project, this opens up a concrete path beyond choosing between the lightweight model and the server model: inject a domain lexicon (product names, internal jargon, acronyms, local proper names) into the language model instead of relying solely on generic vocabulary. It takes more work than a simple download, but it is the lever that comes closest to bringing a Vosk deployment to the quality of a large model in a specific domain, without paying the price in compute or latency.

!
The input format, a source of silent failures
Vosk expects a mono PCM stream, typically sampled at precisely 16 kHz. A stereo file, a different sample rate, or an unconverted compressed format degrades recognition without ever displaying an explicit error—the symptom looks like a bad model, while the real cause is upstream, in the conversion.

#Implementation

The process is simple: install the package (pip install vosk in Python), download a model for the desired language, open an audio stream, feed chunks into the engine, and read the partial and then final results. Bindings exist for Python, Java, Node.js, C#, C++, Rust, and Go, and a WebSocket server lets you connect a web page or mobile app to an engine running on your machine.

Install Vosk and download the lightweight French model
pip install vosk
curl -LO https://alphacephei.com/vosk/models/vosk-model-small-fr-0.22.zip
unzip vosk-model-small-fr-0.22.zip

Two points to keep in mind before going to production: quality depends heavily on the language model you choose—the table above gives you a rough estimate, but test on your own recordings before building on it—and the audio sampling rate must match what the model expects; otherwise, recognition quality degrades without an explicit error message.

#FAQ

Is Vosk free?+
Yes: the alphacep/vosk-api repository is released under the Apache 2.0 license, one of the most permissive open licenses, which covers commercial use without royalties. However, check the license of the specific language model you download: the two official French models are also under Apache 2.0, but some third-party models in the catalog have their own terms.
Do you need a graphics card?+
No, and that’s the whole point: the project is designed to run on a processor, including embedded hardware such as a Raspberry Pi or an Android smartphone, with portable models of around 50 MB—41 MB precisely for the official lightweight French model.
Vosk or Whisper?+
Vosk for live transcription, voice commands, and embedded use, with near-zero latency thanks to its streaming API. Whisper and derivatives such as Faster-Whisper for transcribing files with the best text quality. The two combine very well in the same pipeline.
Does it work in French, and how well?+
Yes, two official models exist: vosk-model-small-fr-0.22 (41 MB, word error rate around 20 to 27 depending on the test) for embedded use, and vosk-model-fr-0.22 (1.4 GB, around 12 to 14) for better server-side accuracy. Alternatives are also available through the LinTO project from the French publisher Linagora.
Can you limit the recognized vocabulary?+
Yes, and it is the setting that most improves accuracy for voice commands: restricting recognition to a list of expected phrases significantly improves reliability, a feature the project explicitly documents as reconfigurable vocabulary, alongside speaker identification.
What model size should you choose for an embedded project?+
The 41 MB lightweight model (vosk-model-small-fr-0.22 for French) remains the reference choice for Raspberry Pi or Android. The 1.4 GB server model nearly halves the word error rate measured on the same test sets, but requires more available memory and compute, ruling it out for a truly embedded deployment.
Can you improve accuracy without changing the model?+
Yes, by customizing the language model with the project's domain vocabulary rather than changing the model size. A recent study built around the Vosk pipeline shows that this approach reduces the word error rate, especially with technical jargon, strong accents, or noisy audio—without the computational cost of a larger model.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.