Faster-Whisper: transcribe quickly, even without GPU
Faster-Whisper (SYSTRAN/faster-whisper, MIT license) is a reimplementation of Whisper built on the CTranslate2 engine. The project claims up to 4 times the speed of openai/whisper, at equal accuracy, while using less memory. In its own CPU benchmark (Intel i7-12700K, small model), it transcribes 13 minutes of audio in 1 min 42 in int8, compared with 6 min 58 for openai/whisper — nearly 4 times faster on a processor, without a graphics card.
Faster-Whisper is a reimplementation of Whisper that transcribes significantly faster and uses substantially less memory, with equivalent text quality. It is not a new model: it uses the same weights, executed by CTranslate2, a different inference engine published by SYSTRAN under the MIT license. It is the default choice for local transcription, and its own benchmark precisely quantifies what “finally makes CPU transcription practical” means in practice.
#What it is—and what it isn’t
Whisper refers both to a family of models and to OpenAI's reference implementation that runs them. Faster-Whisper keeps the former and replaces the latter with CTranslate2, an inference engine optimized for transformer-based models, originally developed for machine translation at OpenNMT. The generated text is the same at equivalent quality; what changes is the time and memory required to produce it, not the content of the final result.
Practical consequence: everything you already know about choosing the size of a Whisper model remains fully applicable, and most local transcription tools you encounter every day already use it under the hood, often without explicitly saying so or documenting it. Another useful detail: unlike openai-whisper, FFmpeg does not need to be installed separately on the system—the audio is decoded by the PyAV library, which bundles its own FFmpeg libraries.
#Where the gain comes from
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
- A specialized inference engine
- CTranslate2, designed to run transformers efficiently, rather than a general-purpose framework such as PyTorch used directly.
- Quantization
- Running the model in 8-bit mode halves its memory footprint and speeds it up, with a generally imperceptible loss in transcription quality.
- Batch processing
- Long files can be split up and processed in parallel (batch_size) instead of strictly one after another, at the cost of higher memory usage.
- Voice activity detection
- Skipping silence prevents the model from working on empty audio—on meeting recordings, that is far from negligible.
#The publisher's quantified gain
SYSTRAN publishes its own benchmark instead of making readers guess: a 13-minute audio transcription compared across several implementations on the same hardware, with the time and memory used by each. Two sets of measurements matter especially for use without a dedicated graphics card—the setup that applies to most enterprise workstations.
| Implementation | Accuracy | Time | RAM |
|---|---|---|---|
| openai/whisper | fp32 | 6 min 58 s | 2,335 MB |
| faster-whisper | fp32 | 2 min 37 sec | 2,257 MB |
| faster-whisper | int8 | 1 min 42 sec | 1 477 MB |
| faster-whisper (batches of 8) | int8 | 51 s | 3,608 MB |
On CPU, the int8 version of Faster-Whisper therefore transcribes 4.1 times faster than openai/whisper, while using less memory (1,477 MB versus 2,335 MB)—consistent with the project's general claim, “up to 4 times faster … while using less memory.” On GPU, the same benchmark with the large-v2 model (RTX 3070 Ti) takes 1 min 03 in fp16 and 59 seconds in int8, versus 2 min 23 for openai/whisper—quantization provides a more modest gain there than on CPU, but video memory usage drops from 4,708 MB to 2,926 MB.
The repository also publishes a comparison with the distil-whisper-large-v3 variant, a lightweight model created through distillation rather than quantization. On a GPU, using the transformers library in fp16 with a batch size of 16, this model takes 46 minutes 12 seconds to transcribe the test corpus, with a word error rate of 14.801; the same setup with Faster-Whisper takes 25 minutes 50 seconds, with a word error rate of 13.527—faster AND more accurate on this specific test, showing that CTranslate2's gain isn't limited to quantization: the runtime engine itself matters, independently of the exact model it runs.
The benchmark also compares Faster-Whisper with whisper.cpp, another widely used reimplementation, especially on Mac and embedded hardware. On the large-v2 model in fp16, whisper.cpp with Flash Attention takes 1 minute 05 seconds and 4,127 MB of video memory, nearly matching Faster-Whisper (1 minute 03 seconds, 4,525 MB)—a useful reminder that Faster-Whisper is not the only fast option, just the one this guide documents in the most detail because its Python ecosystem integrates most easily into a RAG pipeline or an automated meeting-minutes workflow.
#Install it and transcribe a first file
Installation takes one line, with no system dependency to manage for audio decoding thanks to PyAV, which is included directly in the Python package. The minimal code uses three parameters that make all the difference compared with a naive use of the model: explicit device selection (cpu or cuda), the compute type (compute_type, where int8 is a reasonable starting point on CPU), and enabling the voice activity detection filter.
Three lines of code are enough to apply the two settings that explain most of the difference measured in the official benchmark: compute_type="int8" for quantization, and vad_filter=True for voice activity detection. The language="fr" parameter disables automatic language detection, which is most often wrong during the first few seconds of a recording.
#Which model, which precision
| Size | Indicative memory usage at 8-bit | What for |
|---|---|---|
| Small (small) | About 1.5 GB in int8 on CPU, measured on the official benchmark | Quick note-taking, clean audio, one language only |
| Medium | Around 1 to 2 GB | The usual compromise for long recordings |
| Large (large-v2) | About 2.9 GB in int8 on GPU, measured on the official test bench | Polished French, specialized vocabulary, difficult audio |
For French, the difference between an intermediate model and a large model is most noticeable with proper names, acronyms, and passages where several people speak at once. If the transcription then feeds a language model that writes a structured summary, an isolated error on a rare word matters less than intuition suggests; if it must be read as-is by a human, the large model is justified, at the cost of the time and memory specified earlier in this guide.
#The settings that really matter
- 01Force the languageAutomatic detection gets the first few seconds wrong, especially if the recording starts with silence or a short greeting. Specifying the language eliminates an entire class of errors.
- 02Enable voice activity detectionIt avoids transcribing silences and reduces end-of-file hallucinations, where the model invents text from empty audio.
- 03Choose int8 quantizationThis is the setting that explains most of the difference measured on the official benchmark: on CPU, it reduces the time from 2 min 37 to 1 min 42 on the small model, while using less memory.
- 04Enable batch processing if memory allowsbatch_size=8 still halves the time in the official benchmark, at the cost of higher memory usage: reserve it for machines with headroom.
#Three practical uses on an ordinary machine
The official benchmark figures make sense when applied to real-world cases, on a workstation without a dedicated graphics card—the most common situation in business.
- Meeting minutes
- Transcribing a one-hour meeting in int8 on a desktop processor takes roughly seven to eight minutes, extrapolating from figures measured over 13 minutes—more than compatible with transcription started during the coffee break after the meeting itself.
- Podcast or webinar archiving
- A batch job launched overnight on a stack of long files directly benefits from the gain measured with batch_size=8, without requiring you to monitor it until morning.
- Retrospective subtitling of a video
- Faster-Whisper produces the video’s raw text; word-level alignment with WhisperX then adds the precise timestamps needed for a subtitle file that is actually usable in editing.
In all three cases, the common factor is the absence of a graphics card from the equation: CPU benchmarks from the official test bench show that a recent desktop processor is more than sufficient for everyday use, shifting the real question toward model size and quantization rather than buying additional hardware for a simple transcription need.
#Faster-Whisper or something else
| Need | Choice |
|---|---|
| Text, fast, locally, even without a GPU | Faster-Whisper |
| Word-by-word timestamps for subtitling | A pipeline with forced alignment (WhisperX) |
| Know who said what | A chain with speaker separation |
| Live transcription, continuous stream | A stream-oriented engine, such as Vosk |
| A very modest machine with no GPU | Faster-Whisper in int8, small or medium model |
- WhisperX: word-level timestamps and speakers
- Whisper locally: the basics
- From audio file to report
- Vosk: transcribe continuously instead of to a file
- Source: the official Faster-Whisper repository and its benchmark
- Source: CTranslate2 inference engine
- Source: official openai/whisper repository (benchmark reference)
#FAQ
Is Faster-Whisper less accurate than Whisper?+
Do you need a GPU?+
Why does the model make up sentences about silences?+
What model size do you need for French?+
Can it do live streaming?+
What benefit does int8 quantization really provide?+
Is Faster-Whisper Still Faster Than Other Implementations?+
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.