WhisperX: word-level timestamps and locuteurs
WhisperX is a free software wrapper (BSD-2-Clause license, more than 24,000 GitHub stars) around Whisper that adds two things: forced alignment using a phonetic model for truly accurate word-level timestamps, and speaker diarization via pyannote-audio. It all runs locally, with batched inference advertised by the project as up to 70 times faster than real time with Whisper large-v2—at the cost of loading three models instead of one.
Whisper transcribes very well and provides approximate timestamps. WhisperX adds two things that make all the difference once you move beyond a simple document: forced alignment that provides truly accurate word-level timestamps, and speaker diarization that indicates who spoke when. That is what turns a transcription into usable subtitles and a readable meeting summary.
#What Whisper alone doesn't solve
WhisperX is a free BSD-2-Clause-licensed wrapper that adds two things to Whisper that the model alone does not provide: word-level timestamps obtained through forced alignment with a wav2vec2 model, and speaker labels produced by a pyannote model. You need WhisperX if you're creating subtitles that must be timed accurately or meeting minutes that show who spoke. If you only want text, Whisper alone or faster-whisper is enough. The cost is loading three models instead of one, a Hugging Face access token for speaker diarization, and documented limitations: numbers and amounts are not aligned, overlapping speech is handled poorly, and speaker separation remains imperfect. It all runs locally, on a GPU, CPU, or Mac.
Whisper produces excellent text. Its weaknesses appear as soon as you want anything other than text. According to the WhisperX repository, Whisper's timestamps are provided per utterance rather than per word, and can be off by several seconds: invisible in a document, unacceptable for subtitles, where a subtitle arriving one beat too late is worse than no subtitle at all. And Whisper has no notion of speakers: two hours of a meeting come back as an undifferentiated monologue.
WhisperX addresses both by adding steps rather than replacing the model. This design choice preserves the transcription quality you know. The project, released under the BSD-2-Clause license, now has more than 24,000 stars on GitHub—a sign that the need it addresses, time-stamping and attributing a transcript, goes far beyond subtitling.
#The three steps
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
| Step | Role | What it needs |
|---|---|---|
| Voice activity detection (VAD) | The signal is divided into speech regions before transcription | A speech detector (pyannote by default, silero optionally) |
| Transcription | Audio becomes text via faster-whisper, a fast implementation of Whisper | A Whisper model: its size determines quality and memory usage (the command-line default is small) |
| Forced alignment | Each word is positioned exactly in the audio via wav2vec2 | A phonetic model specific to the language |
| Speaker separation | The audio is split by speaker and the turns are labeled using pyannote-audio | The default pyannote speaker-diarization-community-1 model, which requires accepting terms to access |
Voice detection serves two purposes: the repository says it reduces hallucinations and enables segment grouping without degrading the word error rate. The alignment step is the one people underestimate. It's what makes word-level timestamps trustworthy, using a wav2vec2 model to align each word with its actual position in the audio signal. It depends on the language: the repository specifies that a language-specific wav2vec2 model is required, and that default models exist for English, French, German, Spanish, and Italian, plus many other languages through Hugging Face. For a language not on the list, you have to find and test a phonetic model yourself. Transcription itself benefits from batched inference: the project claims up to 70 times real time with the large-v2 model, without specifying the hardware in its README. Treat this figure as an upper bound to verify on your machine.
#Speaker separation, without illusions
Diarization works by grouping vocal characteristics across the entire recording via pyannote-audio. WhisperX’s default model is now speaker-diarization-community-1, released under the CC-BY-4.0 license, which ingests mono audio at 16 kHz (stereo is converted to mono, and other sampling rates are resampled). The model card claims it performs much better than the old 3.1 pipeline. The WhisperX repository is frank about the limitations: overlapping speech is handled only mediocrely, and the separation remains “far from perfect.” The diarization error rates published by pyannote show this: 11.7% on AISHELL-4, 17.0% on AMI (individual microphones), 19.9% on AMI (distant microphone), 26.7% on CALLHOME (part 2), and 46.8% on Ego4D, depending on the corpus. These are academic corpora, not your recordings, but the order of magnitude shows that review is still necessary. The situation gets worse precisely where meetings become complicated: people talking over one another, someone far from the microphone, and similar-sounding voices.
Two practical notes. You can specify a minimum and maximum number of speakers (the --min_speakers and --max_speakers options), and the repository recommends doing so when you know it: the model no longer has to guess. The labels are also anonymous by design: the system says that there were three voices and which turns belong to each one, not who they are. Linking a label to a name is a manual step—and it is also when the transcription becomes personal data.
#Actual accuracy, not benchmark accuracy
The transcription itself deserves a number, because the usual benchmark paints an overly flattering picture. According to VexaScribe, a commercial transcription service, Whisper large-v3 reaches an approximately 2.7% word error rate on LibriSpeech test-clean, a clean single-speaker audio dataset. On real English audio—meetings, podcasts, calls—the same site puts the rate at around 8 to 12%. These are published ballpark figures from an industry player, valid for English: they say nothing about your French recordings. WhisperX does not improve the transcription itself, which remains Whisper's; it adds the ability to determine exactly where each word occurs and who said it, two pieces of information that this error rate does not capture.
| Need | Tool |
|---|---|
| Plain text, fast | Whisper alone or a fast implementation such as faster-whisper |
| Subtitles that appear at exactly the right time | WhisperX: alignment is the whole point |
| Report showing who said what | WhisperX with speaker separation enabled |
| Live transcription, low latency | A flow-oriented engine; this pipeline is made for files |
#What it requires
- The Whisper model size determines everything
- A larger model produces the best text and requires the most memory; the repository notes that a larger model improves timestamp accuracy at the cost of more GPU memory, while a larger alignment model provides little benefit.
- Three models load, not one
- Memory peaks when transcription, alignment, and separation are all resident. The repository's Python example frees each model after its step (memory collection followed by torch.cuda.empty_cache()), a classic remedy on a small card.
- Without a GPU, including on Mac
- The repository provides the command line for the processor and for macOS: --compute_type int8 --device cpu. It publishes no processor processing times: measure on a sample before scheduling a batch. With a NVIDIA GPU, it requires the CUDA 12.8 toolkit, and the project states that large-v2 uses less than 8 GB of memory with beam_size=5.
- Batch processing helps
- Batching is why the advertised speed is possible: the command-line default is --batch_size 8. Lowering it (the repository cites 4) frees up GPU memory.
#Install and run a transcription
Installation is done via pip; the package requires Python 3.10 to 3.13. By default, the command line writes all available formats (srt, vtt, txt, tsv, json, aud), and the --highlight_words True option highlights each word as it is spoken in SRT and VTT subtitles. The JSON includes speaker labels when diarization is enabled.
- 01Install WhisperXpip install whisperx installs the package; with an NVIDIA GPU, first install the CUDA 12.8 toolkit, according to the repository. The repository adds that you may also need to install ffmpeg, referring to Whisper’s installation instructions.
- 02Transcribe and alignRun a first pass with a medium-sized Whisper model (the default is small), specifying the language if it is known in advance: without it, the language is detected automatically, and the alignment model depends on it.
- 03Enable separation if neededAdd --diarize and the Hugging Face access token obtained after accepting the terms for the pyannote model, and specify the number of speakers if you know it.
- 04Review the output fileOpen the generated SRT or JSON and check a few passages at random, especially speaker turns at moments when several people are speaking almost simultaneously.
- Whisper locally: basic transcription
- Produce a meeting summary locally
- Local voice assistant: the complete pipeline
- Transcribing consultations with Whisper
#What alignment will not be able to timestamp
The repository lists limitations to know before promising perfect subtitles. The first concerns numbers: words whose characters do not appear in the alignment model’s dictionary, such as “2014.” or “£13.60,” cannot be aligned and therefore have no timestamps. In a meeting full of amounts and dates, these words appear in the SRT without their own timing. The second is overlapping speech, which Whisper and WhisperX both handle poorly. The third is language: you need a language-specific wav2vec2 model.
Two differences from the original Whisper also explain text discrepancies. To fit in a single pass per batch, inference runs without Whisper timestamps, which, the repository warns, can create differences from the default output. And the condition_on_prev_text option is disabled by default to reduce hallucinations. Since 2026, the --interleaved_context option restores continuous context between segments, which can help with punctuation and technical vocabulary depending on the project, but it is recent: test it on an excerpt before adopting it.
#Concrete use cases
- Subtitling training videos
- Word-level timestamps prevent subtitles from lagging an entire sentence behind the speaker, an immediately visible and annoying flaw in educational content.
- Multi-speaker meeting minutes
- Speaker separation makes it possible to reconstruct who proposed what, information that a summary without attribution cannot faithfully reproduce.
- Podcasts and long-form interviews
- Batch inference is designed for long recordings: that’s what the authors’ article title, “Time-Accurate Speech Transcription of Long-Form Audio,” announces.
- Audio archives to index
- A word-level timestamped transcript lets a search point directly to the exact moment in the recording, not just to the document as a whole.
- Excerpts for social media
- Identifying the exact point where a memorable line is spoken, word for word, makes it easier to cut a short excerpt without listening through the entire video again.
One thing these cases have in common: none requires real-time transcription. An already recorded file is processed afterward, leaving time for batched inference to do its work without latency constraints. When the need becomes live captioning—a streamed conference or an ongoing call—WhisperX is no longer the right tool, regardless of the available machine's power.
#Transcription is an input, not a deliverable
In almost all real-world cases, transcription is not the goal: it feeds a local model that produces meeting minutes, a decision log, or a summary. This affects the expected quality: a transcription error on a rare word matters less than the structure of speaker turns, because that structure is what lets the model correctly attribute what was said.
In practice, a sensible workflow separates the two jobs: WhisperX produces timestamped text labeled by speaker, then a separate language model rereads that text to extract a structured summary. Combining both in a single call—asking a model to summarize the audio directly—deprives the pipeline of the timestamp precision that WhisperX is specifically designed to provide and makes it harder to verify a specific passage afterward.
Another advantage: transcription happens only once, then can be reused for multiple deliverables—short summary, decision log, subtitles—without going back to the audio. The timestamped, labeled text becomes the source of truth.
- Source: official WhisperX repository on GitHub
- Source: pyannote community-1 speaker separation model
- Source: VexaScribe, word error rate for Whisper large-v3 (commercial service)
- Source: whisperx on PyPI
#FAQ
Is WhisperX free?+
How reliable is speaker separation?+
Do you need a GPU?+
WhisperX or Whisper alone?+
Does it work in French?+
Can it transcribe in real time?+
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.