Whisper + Ollama locally: 100% offline transcription ligne
Transcribing a one-hour meeting and then producing a structured summary is exactly what Otter, Fireflies, and Tactiq sell—by sending your audio to a third party. Whisper (OpenAI, but open source and runnable offline) coupled with an LLM Ollama does the same thing on your machine, with no subscription and no data leak. This guide builds the whisper.cpp → Ollama pipeline end to end: installation, model selection, a Bash script, a Python script, and a real-world example of locally transcribing a 1h meeting.
#Why use a whisper + ollama pipeline for local transcription?
A meeting recording contains client names, figures, strategic decisions, and sometimes HR data. SaaS transcription services store these recordings on their servers and use them (according to their terms of service) to train their own models. For a team that takes confidentiality seriously, that is non-negotiable.
OpenAI's Whisper is freely downloadable and runs offline. whisper.cpp (the C++ port by Georgi Gerganov) runs it efficiently on CPU and GPU, without requiring Python or CUDA. Ollama, for its part, exposes a local LLM on http://localhost:11434. The two fit together step by step: audio → text (Whisper) → structured summary (Ollama). Everything stays on your disk.
#Prerequisites
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
- macOS, Linux, or Windows (WSL2)
- whisper.cpp compiles everywhere. On pure Windows, use WSL2—the native build works but requires more tinkering.
- Ollama installed and operational
- A ollama run qwen3.5:9b must respond. Otherwise, go back to the installation guide for your OS.
- C++ compiler and CMake
- build-essential cmake on Debian/Ubuntu, Xcode Command Line Tools on macOS.
- ffmpeg
- To convert your .m4a/.mp3/.mp4 files to 16 kHz WAV, the only format whisper.cpp accepts as input.
- 8 GB of RAM minimum
- 16 GB is comfortable. The large-v3 model loads ~3 GB into RAM/VRAM. A small LLM such as Qwen 3.5 4B adds another ~3 GB.
- Optional GPU
- A RTX 3060 12 GB transcribes 1h of audio in ~4 min with large-v3. On Apple Silicon, Metal is enabled automatically, and it is just as fast.
#1. Install whisper.cpp
No official package: clone the repository and compile it. It takes 2 minutes and produces a portable binary you can move wherever you want.
The main binary is located at build/bin/whisper-cli (formerly named main before the CMake overhaul). This is the one you will call to transcribe.
Install ffmpeg separately if it isn't already installed:
#2. Choose the right Whisper model (tiny → large-v3)
Whisper comes in five sizes, each with two variants: multilingual (the default) and English-only (the .en suffix). For French, stick with multilingual. The table below summarizes what really changes.
- tiny (75 MB)
- Fast but mediocre in French. Intended for quick tests or Raspberry Pis. WER (error rate) ~30% in spontaneous French.
- base (142 MB)
- Acceptable for clear English, weak on noisy French. A good option if you have very limited resources.
- small (466 MB)
- The first serious tier. WER ~10–12% in clean French. Transcribes 1h of audio in ~3 min on a modest GPU.
- medium (1.5 GB)
- Very good quality/speed tradeoff for meeting-room French. WER ~7–8%. Handles punctuation well.
- large-v3 (3.1 GB)
- State of the art. Nearly perfect in clean French, robust with accents and background noise. Prefer it whenever you have >6 GB of free RAM/VRAM.
To transcribe a meeting in French, the practical recommendation is: medium by default, large-v3 if your GPU can handle it. small is a good fallback on a laptop without a dedicated GPU.
Models arrive in models/ in ggml format (the equivalent of gguf for LLMs). You can download multiple sizes and switch between them as needed.
#3. Transcribe an audio file
whisper.cpp only reads mono 16 kHz WAV files. We use ffmpeg to prepare the audio, then whisper-cli to transcribe it.
- -ar 16000
- Resample at 16 kHz (Whisper's training frequency).
- -ac 1
- Forces mono. Whisper ignores stereo anyway.
- -c:a pcm_s16le
- 16-bit little-endian PCM, the expected uncompressed WAV format.
The result is saved to reunion.txt with timestamped segments. A few useful flags to know:
- -l fr
- Forces French. Without this flag, Whisper detects the language automatically—but sometimes gets short phrases wrong.
- -otxt / -osrt / -ovtt / -ojson
- Output format: plain text, SRT or VTT subtitles, or structured JSON with timestamps. You can combine them.
- -of <prefixe>
- Output filename prefix (without the extension).
- -t 8
- Number of CPU threads. By default, whisper.cpp uses 4—raise this to 8 or 16 on a recent CPU.
- --print-progress
- Displays a progress bar. Useful for long files.
#4. Summarize the transcription with Ollama
Once you have the raw transcription, send it to Ollama through its local HTTP API. The /api/chat endpoint is the most direct option for one-shot use.
For a 1-hour meeting, the transcript can easily reach 10,000 to 15,000 tokens. Qwen 3.5 9B accepts 256k of native context, which is more than enough. For an even better summary, Mistral Small 24B (excellent in French) fits under 16 GB of VRAM.
#5. Complete pipeline: meeting-minutes example (1h)
We put everything together in a Python script that takes an audio file as input and outputs a Markdown report. This is the core of a reusable local whisper ollama transcription pipeline.
In a 1h meeting with a RTX 3060 12 GB and the medium model: ~3 minutes for transcription, ~30 seconds for the Qwen 3.5 9B summary. Total: under 4 minutes, zero network requests. On a MacBook M2 16 GB using CPU + Metal: ~6 minutes total.
#Tips and troubleshooting
- Transcription that hallucinates during silences
- Whisper sometimes invents sentences during long silences (“Subtitles by...”, “Thanks for watching”). Enable VAD: --vad --vad-model models/ggml-silero-v5.1.2.bin (download it using the same script).
- Distorted proper names
- Whisper doesn't know your colleagues' names. Pass them in the initial prompt with --prompt "Marie, Julien, Samir, ACME Corp, ProjetX". They will be spelled correctly.
- Compressed audio (Zoom, Teams)
- Zoom MP4 recordings often contain poor-quality mono audio. Prefer Zoom's “audio only” export (M4A), or record in parallel with OBS if quality is a priority.
- ggml_metal_init error on Mac
- You compiled without Metal, or whisper.cpp cannot find the resources. Recompile from scratch: rm -rf build && cmake -B build && cmake --build build --config Release -j.
- Ollama timeout on large context
- The HTTP client cuts out before completion. Increase timeout=600 (10 min) on the requests side. On the server side, OLLAMA_KEEP_ALIVE=30m keeps the model warm.
- Truncated transcription
- If the summary seems to cut off the audio partway through, num_ctx is too low. Check the token count (wc -w on the .txt × 1.3 ≈ tokens) and set num_ctx to twice that.
- Diarization (who is speaking)
- whisper.cpp does not perform diarization. To distinguish speakers, add pyannote.audio to the pipeline before transcription, then annotate each segment. This is another layer, outside the scope here.
#Go further
Your pipeline transcribes and summarizes locally. Three logical directions depending on your needs:
- Medical report with faster-whisper
- The consultation transcription guide uses faster-whisper (an optimized Python port) with a focus on medical confidentiality, pyannote diarization, and French healthcare GDPR compliance. An interesting variation on the same pattern.
- Choose the right summarization LLM
- Summarization is where quality varies the most. The guide to choosing your quantization compares Q4_K_M, Q5_K_M, and Q8_0—a noticeable difference on long, nuanced text.
- Automate the pipeline with n8n
- To trigger the chain as soon as an audio file arrives (watched folder, email, Slack), the n8n + Ollama automation guide shows you how to build the workflow without code.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.