Intermediate 14 minAudio

WhisperX: word-level timestamps and locuteurs

Direct response

WhisperX is a free software wrapper (BSD-2-Clause license, more than 24,000 GitHub stars) around Whisper that adds two things: forced alignment using a phonetic model for truly accurate word-level timestamps, and speaker diarization via pyannote-audio. It all runs locally, with batched inference advertised by the project as up to 70 times faster than real time with Whisper large-v2—at the cost of loading three models instead of one.

Whisper transcribes very well and provides approximate timestamps. WhisperX adds two things that make all the difference once you move beyond a simple document: forced alignment that provides truly accurate word-level timestamps, and speaker diarization that indicates who spoke when. That is what turns a transcription into usable subtitles and a readable meeting summary.

By Mohamed Meguedmi·Update 2026-09-29·Tested on Windows, macOS, and Linux

#What Whisper alone doesn't solve

WhisperX is a free BSD-2-Clause-licensed wrapper that adds two things to Whisper that the model alone does not provide: word-level timestamps obtained through forced alignment with a wav2vec2 model, and speaker labels produced by a pyannote model. You need WhisperX if you're creating subtitles that must be timed accurately or meeting minutes that show who spoke. If you only want text, Whisper alone or faster-whisper is enough. The cost is loading three models instead of one, a Hugging Face access token for speaker diarization, and documented limitations: numbers and amounts are not aligned, overlapping speech is handled poorly, and speaker separation remains imperfect. It all runs locally, on a GPU, CPU, or Mac.

Whisper produces excellent text. Its weaknesses appear as soon as you want anything other than text. According to the WhisperX repository, Whisper's timestamps are provided per utterance rather than per word, and can be off by several seconds: invisible in a document, unacceptable for subtitles, where a subtitle arriving one beat too late is worse than no subtitle at all. And Whisper has no notion of speakers: two hours of a meeting come back as an undifferentiated monologue.

WhisperX addresses both by adding steps rather than replacing the model. This design choice preserves the transcription quality you know. The project, released under the BSD-2-Clause license, now has more than 24,000 stars on GitHub—a sign that the need it addresses, time-stamping and attributing a transcript, goes far beyond subtitling.

#The three steps

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates
What each step does and what it requires
StepRoleWhat it needs
Voice activity detection (VAD)The signal is divided into speech regions before transcriptionA speech detector (pyannote by default, silero optionally)
TranscriptionAudio becomes text via faster-whisper, a fast implementation of WhisperA Whisper model: its size determines quality and memory usage (the command-line default is small)
Forced alignmentEach word is positioned exactly in the audio via wav2vec2A phonetic model specific to the language
Speaker separationThe audio is split by speaker and the turns are labeled using pyannote-audioThe default pyannote speaker-diarization-community-1 model, which requires accepting terms to access

Voice detection serves two purposes: the repository says it reduces hallucinations and enables segment grouping without degrading the word error rate. The alignment step is the one people underestimate. It's what makes word-level timestamps trustworthy, using a wav2vec2 model to align each word with its actual position in the audio signal. It depends on the language: the repository specifies that a language-specific wav2vec2 model is required, and that default models exist for English, French, German, Spanish, and Italian, plus many other languages through Hugging Face. For a language not on the list, you have to find and test a phonetic model yourself. Transcription itself benefits from batched inference: the project claims up to 70 times real time with the large-v2 model, without specifying the hardware in its README. Treat this figure as an upper bound to verify on your machine.

#Speaker separation, without illusions

Diarization works by grouping vocal characteristics across the entire recording via pyannote-audio. WhisperX’s default model is now speaker-diarization-community-1, released under the CC-BY-4.0 license, which ingests mono audio at 16 kHz (stereo is converted to mono, and other sampling rates are resampled). The model card claims it performs much better than the old 3.1 pipeline. The WhisperX repository is frank about the limitations: overlapping speech is handled only mediocrely, and the separation remains “far from perfect.” The diarization error rates published by pyannote show this: 11.7% on AISHELL-4, 17.0% on AMI (individual microphones), 19.9% on AMI (distant microphone), 26.7% on CALLHOME (part 2), and 46.8% on Ego4D, depending on the corpus. These are academic corpora, not your recordings, but the order of magnitude shows that review is still necessary. The situation gets worse precisely where meetings become complicated: people talking over one another, someone far from the microphone, and similar-sounding voices.

Two practical notes. You can specify a minimum and maximum number of speakers (the --min_speakers and --max_speakers options), and the repository recommends doing so when you know it: the model no longer has to guess. The labels are also anonymous by design: the system says that there were three voices and which turns belong to each one, not who they are. Linking a label to a name is a manual step—and it is also when the transcription becomes personal data.

!
Processing locally does not constitute consent
A transcription with speaker labels is personal data, whether or not it leaves your machine. Local processing is the right technical posture; it does not exempt you from informing the recorded individuals and obtaining their consent.
i
What access requirements mean in practice
The speaker-separation model is hosted on Hugging Face. The repository is public, but the model page requires you to accept its terms to access the files, then create a read-access token to pass with --hf_token. It is free, not a payment, but it is a step you do not want to discover halfway through a batch job.

#Actual accuracy, not benchmark accuracy

The transcription itself deserves a number, because the usual benchmark paints an overly flattering picture. According to VexaScribe, a commercial transcription service, Whisper large-v3 reaches an approximately 2.7% word error rate on LibriSpeech test-clean, a clean single-speaker audio dataset. On real English audio—meetings, podcasts, calls—the same site puts the rate at around 8 to 12%. These are published ballpark figures from an industry player, valid for English: they say nothing about your French recordings. WhisperX does not improve the transcription itself, which remains Whisper's; it adds the ability to determine exactly where each word occurs and who said it, two pieces of information that this error rate does not capture.

Choose the right tool for the job
NeedTool
Plain text, fastWhisper alone or a fast implementation such as faster-whisper
Subtitles that appear at exactly the right timeWhisperX: alignment is the whole point
Report showing who said whatWhisperX with speaker separation enabled
Live transcription, low latencyA flow-oriented engine; this pipeline is made for files

#What it requires

The Whisper model size determines everything
A larger model produces the best text and requires the most memory; the repository notes that a larger model improves timestamp accuracy at the cost of more GPU memory, while a larger alignment model provides little benefit.
Three models load, not one
Memory peaks when transcription, alignment, and separation are all resident. The repository's Python example frees each model after its step (memory collection followed by torch.cuda.empty_cache()), a classic remedy on a small card.
Without a GPU, including on Mac
The repository provides the command line for the processor and for macOS: --compute_type int8 --device cpu. It publishes no processor processing times: measure on a sample before scheduling a batch. With a NVIDIA GPU, it requires the CUDA 12.8 toolkit, and the project states that large-v2 uses less than 8 GB of memory with beam_size=5.
Batch processing helps
Batching is why the advertised speed is possible: the command-line default is --batch_size 8. Lowering it (the repository cites 4) frees up GPU memory.
→
Running out of GPU memory: three levers
The repository lists three ways to reduce memory: lower --batch_size (for example, to 4), choose a smaller model (--model base), or use a lighter computation type (--compute_type int8). It notes that the last two can reduce quality. So start with the batch size.

#Install and run a transcription

Installation is done via pip; the package requires Python 3.10 to 3.13. By default, the command line writes all available formats (srt, vtt, txt, tsv, json, aud), and the --highlight_words True option highlights each word as it is spoken in SRT and VTT subtitles. The JSON includes speaker labels when diarization is enabled.

  1. 01
    Install WhisperX
    pip install whisperx installs the package; with an NVIDIA GPU, first install the CUDA 12.8 toolkit, according to the repository. The repository adds that you may also need to install ffmpeg, referring to Whisper’s installation instructions.
  2. 02
    Transcribe and align
    Run a first pass with a medium-sized Whisper model (the default is small), specifying the language if it is known in advance: without it, the language is detected automatically, and the alignment model depends on it.
  3. 03
    Enable separation if needed
    Add --diarize and the Hugging Face access token obtained after accepting the terms for the pyannote model, and specify the number of speakers if you know it.
  4. 04
    Review the output file
    Open the generated SRT or JSON and check a few passages at random, especially speaker turns at moments when several people are speaking almost simultaneously.
Terminal
pip install whisperx

# Avec GPU : transcription, alignement et séparation des locuteurs
whisperx reunion.mp3 --model medium --language fr \
  --diarize --hf_token VOTRE_JETON --min_speakers 2 --max_speakers 5

# Sur processeur ou sur Mac
whisperx reunion.mp3 --model medium --language fr --compute_type int8 --device cpu

#What alignment will not be able to timestamp

The repository lists limitations to know before promising perfect subtitles. The first concerns numbers: words whose characters do not appear in the alignment model’s dictionary, such as “2014.” or “£13.60,” cannot be aligned and therefore have no timestamps. In a meeting full of amounts and dates, these words appear in the SRT without their own timing. The second is overlapping speech, which Whisper and WhisperX both handle poorly. The third is language: you need a language-specific wav2vec2 model.

Two differences from the original Whisper also explain text discrepancies. To fit in a single pass per batch, inference runs without Whisper timestamps, which, the repository warns, can create differences from the default output. And the condition_on_prev_text option is disabled by default to reduce hallucinations. Since 2026, the --interleaved_context option restores continuous context between segments, which can help with punctuation and technical vocabulary depending on the project, but it is recent: test it on an excerpt before adopting it.

#Concrete use cases

Subtitling training videos
Word-level timestamps prevent subtitles from lagging an entire sentence behind the speaker, an immediately visible and annoying flaw in educational content.
Multi-speaker meeting minutes
Speaker separation makes it possible to reconstruct who proposed what, information that a summary without attribution cannot faithfully reproduce.
Podcasts and long-form interviews
Batch inference is designed for long recordings: that’s what the authors’ article title, “Time-Accurate Speech Transcription of Long-Form Audio,” announces.
Audio archives to index
A word-level timestamped transcript lets a search point directly to the exact moment in the recording, not just to the document as a whole.
Excerpts for social media
Identifying the exact point where a memorable line is spoken, word for word, makes it easier to cut a short excerpt without listening through the entire video again.

One thing these cases have in common: none requires real-time transcription. An already recorded file is processed afterward, leaving time for batched inference to do its work without latency constraints. When the need becomes live captioning—a streamed conference or an ongoing call—WhisperX is no longer the right tool, regardless of the available machine's power.

#Transcription is an input, not a deliverable

In almost all real-world cases, transcription is not the goal: it feeds a local model that produces meeting minutes, a decision log, or a summary. This affects the expected quality: a transcription error on a rare word matters less than the structure of speaker turns, because that structure is what lets the model correctly attribute what was said.

In practice, a sensible workflow separates the two jobs: WhisperX produces timestamped text labeled by speaker, then a separate language model rereads that text to extract a structured summary. Combining both in a single call—asking a model to summarize the audio directly—deprives the pipeline of the timestamp precision that WhisperX is specifically designed to provide and makes it harder to verify a specific passage afterward.

Another advantage: transcription happens only once, then can be reused for multiple deliverables—short summary, decision log, subtitles—without going back to the audio. The timestamped, labeled text becomes the source of truth.

#FAQ

Is WhisperX free?+
Yes, it is a free project under the BSD-2-Clause license that runs locally and has more than 24,000 stars on GitHub. Speaker separation uses a pyannote model hosted on Hugging Face; you must accept its terms and create a token: it’s free, with no payment required, but you need to do this before the first batch processing job.
How reliable is speaker separation?+
Accurate on clean audio with distinct voices, but more fragile when voices overlap: the repository acknowledges that separation is « far from perfect ». Errors reported by pyannote range from 11.7% to 46.8% depending on the corpus. Specifying the minimum and maximum number of speakers helps, and human review is still necessary.
Do you need a GPU?+
No: the repository provides the command for the processor and for macOS (--compute_type int8 --device cpu), but publishes no processing time for them, so measure first. With a NVIDIA GPU, the project reports less than 8 GB of memory for large-v2 with beam_size=5, and memory peaks when transcription, alignment, and diarization are loaded together.
WhisperX or Whisper alone?+
Whisper alone if you only want text and do not need precise synchronization, for example to index the content of an interview. WhisperX as soon as you need word-level timestamps for subtitles that appear at the right time, or speaker labels to transcribe a multi-speaker meeting in a readable, usable format.
Does it work in French?+
Yes: the repository lists French among the five languages (along with English, German, Spanish, and Italian) that have a default alignment model. Pass --language fr. For a language absent from this list, you must find a wav2vec2 phonetic model yourself and test it before relying on accurate word-level timestamps.
Can it transcribe in real time?+
No, it is designed for files that have already been recorded, not for a live stream. Live captioning requires a stream-oriented engine, with a different latency–accuracy tradeoff from the one sought here, and the batched inference that makes WhisperX fast on an entire file no longer makes sense for a continuous audio stream arriving piece by piece.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.