Beginner 11 minAudio

Faster-Whisper: transcribe quickly, even without GPU

Direct response

Faster-Whisper (SYSTRAN/faster-whisper, MIT license) is a reimplementation of Whisper built on the CTranslate2 engine. The project claims up to 4 times the speed of openai/whisper, at equal accuracy, while using less memory. In its own CPU benchmark (Intel i7-12700K, small model), it transcribes 13 minutes of audio in 1 min 42 in int8, compared with 6 min 58 for openai/whisper — nearly 4 times faster on a processor, without a graphics card.

Faster-Whisper is a reimplementation of Whisper that transcribes significantly faster and uses substantially less memory, with equivalent text quality. It is not a new model: it uses the same weights, executed by CTranslate2, a different inference engine published by SYSTRAN under the MIT license. It is the default choice for local transcription, and its own benchmark precisely quantifies what “finally makes CPU transcription practical” means in practice.

By Mohamed Meguedmi·Update 2026-09-28·Tested on Windows, macOS, and Linux

#What it is—and what it isn’t

Whisper refers both to a family of models and to OpenAI's reference implementation that runs them. Faster-Whisper keeps the former and replaces the latter with CTranslate2, an inference engine optimized for transformer-based models, originally developed for machine translation at OpenNMT. The generated text is the same at equivalent quality; what changes is the time and memory required to produce it, not the content of the final result.

Practical consequence: everything you already know about choosing the size of a Whisper model remains fully applicable, and most local transcription tools you encounter every day already use it under the hood, often without explicitly saying so or documenting it. Another useful detail: unlike openai-whisper, FFmpeg does not need to be installed separately on the system—the audio is decoded by the PyAV library, which bundles its own FFmpeg libraries.

#Where the gain comes from

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates
A specialized inference engine
CTranslate2, designed to run transformers efficiently, rather than a general-purpose framework such as PyTorch used directly.
Quantization
Running the model in 8-bit mode halves its memory footprint and speeds it up, with a generally imperceptible loss in transcription quality.
Batch processing
Long files can be split up and processed in parallel (batch_size) instead of strictly one after another, at the cost of higher memory usage.
Voice activity detection
Skipping silence prevents the model from working on empty audio—on meeting recordings, that is far from negligible.

#The publisher's quantified gain

SYSTRAN publishes its own benchmark instead of making readers guess: a 13-minute audio transcription compared across several implementations on the same hardware, with the time and memory used by each. Two sets of measurements matter especially for use without a dedicated graphics card—the setup that applies to most enterprise workstations.

Small model on CPU (Intel Core i7-12700K, 8 threads) — source: official repository benchmark
ImplementationAccuracyTimeRAM
openai/whisperfp326 min 58 s2,335 MB
faster-whisperfp322 min 37 sec2,257 MB
faster-whisperint81 min 42 sec1 477 MB
faster-whisper (batches of 8)int851 s3,608 MB

On CPU, the int8 version of Faster-Whisper therefore transcribes 4.1 times faster than openai/whisper, while using less memory (1,477 MB versus 2,335 MB)—consistent with the project's general claim, “up to 4 times faster … while using less memory.” On GPU, the same benchmark with the large-v2 model (RTX 3070 Ti) takes 1 min 03 in fp16 and 59 seconds in int8, versus 2 min 23 for openai/whisper—quantization provides a more modest gain there than on CPU, but video memory usage drops from 4,708 MB to 2,926 MB.

The repository also publishes a comparison with the distil-whisper-large-v3 variant, a lightweight model created through distillation rather than quantization. On a GPU, using the transformers library in fp16 with a batch size of 16, this model takes 46 minutes 12 seconds to transcribe the test corpus, with a word error rate of 14.801; the same setup with Faster-Whisper takes 25 minutes 50 seconds, with a word error rate of 13.527—faster AND more accurate on this specific test, showing that CTranslate2's gain isn't limited to quantization: the runtime engine itself matters, independently of the exact model it runs.

The benchmark also compares Faster-Whisper with whisper.cpp, another widely used reimplementation, especially on Mac and embedded hardware. On the large-v2 model in fp16, whisper.cpp with Flash Attention takes 1 minute 05 seconds and 4,127 MB of video memory, nearly matching Faster-Whisper (1 minute 03 seconds, 4,525 MB)—a useful reminder that Faster-Whisper is not the only fast option, just the one this guide documents in the most detail because its Python ecosystem integrates most easily into a RAG pipeline or an automated meeting-minutes workflow.

→
Batch processing changes the equation
With batch_size=8, the small model in int8 drops from 1 minute 42 seconds to 51 seconds on the same processor, at the cost of memory rising to 3,608 MB. On a machine with available RAM, this is often the most cost-effective setting before you even consider a GPU.

#Install it and transcribe a first file

Installation takes one line, with no system dependency to manage for audio decoding thanks to PyAV, which is included directly in the Python package. The minimal code uses three parameters that make all the difference compared with a naive use of the model: explicit device selection (cpu or cuda), the compute type (compute_type, where int8 is a reasonable starting point on CPU), and enabling the voice activity detection filter.

Install Faster-Whisper
pip install faster-whisper
Transcribe a file in French, on a CPU
from faster_whisper import WhisperModel

model = WhisperModel("small", device="cpu", compute_type="int8")
segments, info = model.transcribe("reunion.wav", language="fr", vad_filter=True)

for segment in segments:
    print(f"[{segment.start:.1f}s -> {segment.end:.1f}s] {segment.text}")

Three lines of code are enough to apply the two settings that explain most of the difference measured in the official benchmark: compute_type="int8" for quantization, and vad_filter=True for voice activity detection. The language="fr" parameter disables automatic language detection, which is most often wrong during the first few seconds of a recording.

#Which model, which precision

Choose the size based on usage
SizeIndicative memory usage at 8-bitWhat for
Small (small)About 1.5 GB in int8 on CPU, measured on the official benchmarkQuick note-taking, clean audio, one language only
MediumAround 1 to 2 GBThe usual compromise for long recordings
Large (large-v2)About 2.9 GB in int8 on GPU, measured on the official test benchPolished French, specialized vocabulary, difficult audio

For French, the difference between an intermediate model and a large model is most noticeable with proper names, acronyms, and passages where several people speak at once. If the transcription then feeds a language model that writes a structured summary, an isolated error on a rare word matters less than intuition suggests; if it must be read as-is by a human, the large model is justified, at the cost of the time and memory specified earlier in this guide.

#The settings that really matter

  1. 01
    Force the language
    Automatic detection gets the first few seconds wrong, especially if the recording starts with silence or a short greeting. Specifying the language eliminates an entire class of errors.
  2. 02
    Enable voice activity detection
    It avoids transcribing silences and reduces end-of-file hallucinations, where the model invents text from empty audio.
  3. 03
    Choose int8 quantization
    This is the setting that explains most of the difference measured on the official benchmark: on CPU, it reduces the time from 2 min 37 to 1 min 42 on the small model, while using less memory.
  4. 04
    Enable batch processing if memory allows
    batch_size=8 still halves the time in the official benchmark, at the cost of higher memory usage: reserve it for machines with headroom.
i
The hallucinations of silence
Whisper has a known quirk: in a segment with no speech, it may produce a recurring phrase—a polite formula or a title sequence. Voice activity detection is the primary remedy; a quick check of repeated segments complements it.

#Three practical uses on an ordinary machine

The official benchmark figures make sense when applied to real-world cases, on a workstation without a dedicated graphics card—the most common situation in business.

Meeting minutes
Transcribing a one-hour meeting in int8 on a desktop processor takes roughly seven to eight minutes, extrapolating from figures measured over 13 minutes—more than compatible with transcription started during the coffee break after the meeting itself.
Podcast or webinar archiving
A batch job launched overnight on a stack of long files directly benefits from the gain measured with batch_size=8, without requiring you to monitor it until morning.
Retrospective subtitling of a video
Faster-Whisper produces the video’s raw text; word-level alignment with WhisperX then adds the precise timestamps needed for a subtitle file that is actually usable in editing.

In all three cases, the common factor is the absence of a graphics card from the equation: CPU benchmarks from the official test bench show that a recent desktop processor is more than sufficient for everyday use, shifting the real question toward model size and quantization rather than buying additional hardware for a simple transcription need.

#Faster-Whisper or something else

The tool for each need
NeedChoice
Text, fast, locally, even without a GPUFaster-Whisper
Word-by-word timestamps for subtitlingA pipeline with forced alignment (WhisperX)
Know who said whatA chain with speaker separation
Live transcription, continuous streamA stream-oriented engine, such as Vosk
A very modest machine with no GPUFaster-Whisper in int8, small or medium model

#FAQ

Is Faster-Whisper less accurate than Whisper?+
No, at the same model size, quality is equivalent: these are exactly the same weights run by CTranslate2, a different engine from openai/whisper. The project itself reports identical accuracy, and 8-bit quantization generally has an entirely imperceptible effect on the resulting transcription.
Do you need a GPU?+
No, and the figures demonstrate why: on an Intel Core i7-12700K, the small model in int8 transcribes 13 minutes of audio in 1 min 42 sec, compared with 6 min 58 sec for openai/whisper on the same processor. A GPU makes it even faster, but isn’t required for reasonable use.
Why does the model make up sentences about silences?+
This is a known behavior of Whisper on passages without speech, inherited from the original model and therefore also present in Faster-Whisper because they use the same weights. Enabling voice activity detection is the primary remedy, followed by a targeted review of repeated passages that reveal the problem in a long recording.
What model size do you need for French?+
The medium model is sufficient if the text will then feed an automatic summary, where an isolated error on a rare word matters little. For a transcript intended to be read as is, the large model (large-v2) is worthwhile for proper names and acronyms, at the cost of about 2.9 GB of int8 memory according to the official benchmark.
Can it do live streaming?+
It is designed to process files that have already been recorded, not a continuous real-time stream. Live captioning requires a stream-oriented engine such as Vosk, with a different latency-versus-accuracy tradeoff—the two tools work well together in the same pipeline, one for immediate feedback and the other for the final version.
What benefit does int8 quantization really provide?+
On the official test bench, it reduces transcription time for the small model from 2 min 37 (fp32) to 1 min 42 (int8) on the CPU, while memory drops from 2 257 MB to 1 477 MB. This isolated setting has the greatest impact before considering a GPU or batch processing.
Is Faster-Whisper Still Faster Than Other Implementations?+
Not systematically: on the large-v2 model in fp16 with beam size 5, whisper.cpp with Flash Attention takes 1 min 05 versus 1 min 03 for Faster-Whisper, essentially tied. The gap widens mainly against the official openai/whisper implementation and transformers, where Faster-Whisper remains clearly ahead in the published benchmark.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.