Intermediate 18 minAudio

Whisper + Ollama locally: 100% offline transcription ligne

Transcribing a one-hour meeting and then producing a structured summary is exactly what Otter, Fireflies, and Tactiq sell—by sending your audio to a third party. Whisper (OpenAI, but open source and runnable offline) coupled with an LLM Ollama does the same thing on your machine, with no subscription and no data leak. This guide builds the whisper.cpp → Ollama pipeline end to end: installation, model selection, a Bash script, a Python script, and a real-world example of locally transcribing a 1h meeting.

By Mohamed Meguedmi·Update 2026-08-27·Tested on Windows, macOS, and Linux

#Why use a whisper + ollama pipeline for local transcription?

A meeting recording contains client names, figures, strategic decisions, and sometimes HR data. SaaS transcription services store these recordings on their servers and use them (according to their terms of service) to train their own models. For a team that takes confidentiality seriously, that is non-negotiable.

OpenAI's Whisper is freely downloadable and runs offline. whisper.cpp (the C++ port by Georgi Gerganov) runs it efficiently on CPU and GPU, without requiring Python or CUDA. Ollama, for its part, exposes a local LLM on http://localhost:11434. The two fit together step by step: audio → text (Whisper) → structured summary (Ollama). Everything stays on your disk.

i
The breakdown
Whisper ≠ LLM. Whisper is an automatic speech recognition (ASR) model: it converts audio into raw text. Summarization, formatting, and decision extraction are the work of an LLM behind it. These two models have nothing to do with each other and do not replace one another.

#Prerequisites

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates
macOS, Linux, or Windows (WSL2)
whisper.cpp compiles everywhere. On pure Windows, use WSL2—the native build works but requires more tinkering.
Ollama installed and operational
A ollama run qwen3.5:9b must respond. Otherwise, go back to the installation guide for your OS.
C++ compiler and CMake
build-essential cmake on Debian/Ubuntu, Xcode Command Line Tools on macOS.
ffmpeg
To convert your .m4a/.mp3/.mp4 files to 16 kHz WAV, the only format whisper.cpp accepts as input.
8 GB of RAM minimum
16 GB is comfortable. The large-v3 model loads ~3 GB into RAM/VRAM. A small LLM such as Qwen 3.5 4B adds another ~3 GB.
Optional GPU
A RTX 3060 12 GB transcribes 1h of audio in ~4 min with large-v3. On Apple Silicon, Metal is enabled automatically, and it is just as fast.

#1. Install whisper.cpp

No official package: clone the repository and compile it. It takes 2 minutes and produces a portable binary you can move wherever you want.

Clone + CPU build
git clone https://github.com/ggerganov/whisper.cpp.git
cd whisper.cpp
cmake -B build
cmake --build build --config Release -j

The main binary is located at build/bin/whisper-cli (formerly named main before the CMake overhaul). This is the one you will call to transcribe.

→
GPU acceleration
For CUDA (NVIDIA), add -DGGML_CUDA=1 to the cmake -B build command. For Apple Silicon, Metal is enabled by default — nothing to configure. For AMD (ROCm): -DGGML_HIPBLAS=1. The binary then detects the GPU at launch.
CUDA variant
cmake -B build -DGGML_CUDA=1
cmake --build build --config Release -j

Install ffmpeg separately if it isn't already installed:

ffmpeg
# Debian / Ubuntu / WSL2
sudo apt install ffmpeg

# macOS
brew install ffmpeg

#2. Choose the right Whisper model (tiny → large-v3)

Whisper comes in five sizes, each with two variants: multilingual (the default) and English-only (the .en suffix). For French, stick with multilingual. The table below summarizes what really changes.

tiny (75 MB)
Fast but mediocre in French. Intended for quick tests or Raspberry Pis. WER (error rate) ~30% in spontaneous French.
base (142 MB)
Acceptable for clear English, weak on noisy French. A good option if you have very limited resources.
small (466 MB)
The first serious tier. WER ~10–12% in clean French. Transcribes 1h of audio in ~3 min on a modest GPU.
medium (1.5 GB)
Very good quality/speed tradeoff for meeting-room French. WER ~7–8%. Handles punctuation well.
large-v3 (3.1 GB)
State of the art. Nearly perfect in clean French, robust with accents and background noise. Prefer it whenever you have >6 GB of free RAM/VRAM.

To transcribe a meeting in French, the practical recommendation is: medium by default, large-v3 if your GPU can handle it. small is a good fallback on a laptop without a dedicated GPU.

Download a model
# Depuis le dossier whisper.cpp
./models/download-ggml-model.sh medium
# ou
./models/download-ggml-model.sh large-v3

Models arrive in models/ in ggml format (the equivalent of gguf for LLMs). You can download multiple sizes and switch between them as needed.

i
Quantized versions
If VRAM is tight, the script accepts quantized variants: large-v3-q5_0 (≈1.1 GB) or medium-q5_0 (≈540 MB). The quality loss is marginal in French, while the memory savings are real. Test them on your own audio.

#3. Transcribe an audio file

whisper.cpp only reads mono 16 kHz WAV files. We use ffmpeg to prepare the audio, then whisper-cli to transcribe it.

Conversion to 16 kHz WAV
ffmpeg -i reunion.m4a -ar 16000 -ac 1 -c:a pcm_s16le reunion.wav
-ar 16000
Resample at 16 kHz (Whisper's training frequency).
-ac 1
Forces mono. Whisper ignores stereo anyway.
-c:a pcm_s16le
16-bit little-endian PCM, the expected uncompressed WAV format.
Transcription
./build/bin/whisper-cli \
  -m models/ggml-medium.bin \
  -f reunion.wav \
  -l fr \
  -otxt \
  -of reunion

The result is saved to reunion.txt with timestamped segments. A few useful flags to know:

-l fr
Forces French. Without this flag, Whisper detects the language automatically—but sometimes gets short phrases wrong.
-otxt / -osrt / -ovtt / -ojson
Output format: plain text, SRT or VTT subtitles, or structured JSON with timestamps. You can combine them.
-of <prefixe>
Output filename prefix (without the extension).
-t 8
Number of CPU threads. By default, whisper.cpp uses 4—raise this to 8 or 16 on a recent CPU.
--print-progress
Displays a progress bar. Useful for long files.
→
JSON format for the pipeline
To pass the output to an LLM, prefer -ojson: you get the segments with their start/end timestamps. This lets you cite a precise timestamp in the final summary ("decision made at 23:14").

#4. Summarize the transcription with Ollama

Once you have the raw transcription, send it to Ollama through its local HTTP API. The /api/chat endpoint is the most direct option for one-shot use.

Prerequisites Ollama
ollama pull qwen3.5:9b
# Ou pour un résumé plus fin d'une longue réunion :
ollama pull mistral-small

For a 1-hour meeting, the transcript can easily reach 10,000 to 15,000 tokens. Qwen 3.5 9B accepts 256k of native context, which is more than enough. For an even better summary, Mistral Small 24B (excellent in French) fits under 16 GB of VRAM.

Quick test with curl
curl -s http://localhost:11434/api/chat -d '{
  "model": "qwen3.5:9b",
  "stream": false,
  "options": {"temperature": 0.2, "num_ctx": 16384},
  "messages": [
    {"role": "system", "content": "Tu résumes une réunion de travail en français."},
    {"role": "user", "content": "<COLLE ICI LE CONTENU DE reunion.txt>"}
  ]
}' | jq -r '.message.content'
!
num_ctx, not by default
Ollama loads its models with a context of 2048 tokens by default, which would silently truncate your transcription. Explicitly pass num_ctx (8192, 16384, 32768…) in options on every call — or lock it in via a Modelfile.

#5. Complete pipeline: meeting-minutes example (1h)

We put everything together in a Python script that takes an audio file as input and outputs a Markdown report. This is the core of a reusable local whisper ollama transcription pipeline.

Python dependencies
pip install requests
pipeline_reunion.py
#!/usr/bin/env python3
"""Audio -> transcription Whisper -> compte-rendu via Ollama."""
import subprocess, sys, requests, pathlib, json

WHISPER_BIN = "./build/bin/whisper-cli"
WHISPER_MODEL = "models/ggml-medium.bin"
OLLAMA_URL = "http://localhost:11434/api/chat"
LLM_MODEL = "qwen3.5:9b"

SYSTEM_CR = """Tu es un assistant qui produit des comptes-rendus de
réunion en français à partir d'une transcription brute.

FORMAT DE SORTIE (Markdown) :
# Compte-rendu
## Participants identifiables
## Sujets abordés (1 paragraphe par sujet)
## Décisions prises
## Actions à mener (avec porteur si mentionné)
## Points en suspens

RÈGLES :
- N'invente AUCUN nom, AUCUNE décision, AUCUNE action.
- Si une info manque, écris \"Non précisé\".
- Sois concis : un compte-rendu se lit en 2 minutes.
"""

def to_wav(src: pathlib.Path) -> pathlib.Path:
    dst = src.with_suffix(".wav")
    subprocess.run([
        "ffmpeg", "-y", "-i", str(src),
        "-ar", "16000", "-ac", "1", "-c:a", "pcm_s16le",
        str(dst),
    ], check=True)
    return dst

def transcribe(wav: pathlib.Path) -> str:
    prefix = wav.with_suffix("")
    subprocess.run([
        WHISPER_BIN, "-m", WHISPER_MODEL, "-f", str(wav),
        "-l", "fr", "-otxt", "-of", str(prefix),
        "--print-progress",
    ], check=True)
    return prefix.with_suffix(".txt").read_text(encoding="utf-8")

def summarize(transcription: str) -> str:
    r = requests.post(OLLAMA_URL, json={
        "model": LLM_MODEL,
        "stream": False,
        "options": {"temperature": 0.2, "num_ctx": 16384},
        "messages": [
            {"role": "system", "content": SYSTEM_CR},
            {"role": "user",   "content": transcription},
        ],
    }, timeout=600)
    r.raise_for_status()
    return r.json()["message"]["content"]

if __name__ == "__main__":
    src = pathlib.Path(sys.argv[1])
    wav = to_wav(src)
    print(f"[1/3] Audio converti -> {wav}")
    texte = transcribe(wav)
    print(f"[2/3] Transcription : {len(texte)} caractères")
    cr = summarize(texte)
    out = src.with_name(src.stem + "_compte-rendu.md")
    out.write_text(cr, encoding="utf-8")
    print(f"[3/3] Compte-rendu : {out}")
Launch
python pipeline_reunion.py reunion.m4a

In a 1h meeting with a RTX 3060 12 GB and the medium model: ~3 minutes for transcription, ~30 seconds for the Qwen 3.5 9B summary. Total: under 4 minutes, zero network requests. On a MacBook M2 16 GB using CPU + Metal: ~6 minutes total.

→
Very long meetings
Beyond 90 minutes or 30,000 transcription tokens, split the text into 10,000-token chunks, summarize each chunk, then ask the LLM to merge the summaries. This is the classic map-reduce pattern. Otherwise, summary quality collapses, even if num_ctx accepts it.

#Tips and troubleshooting

Transcription that hallucinates during silences
Whisper sometimes invents sentences during long silences (“Subtitles by...”, “Thanks for watching”). Enable VAD: --vad --vad-model models/ggml-silero-v5.1.2.bin (download it using the same script).
Distorted proper names
Whisper doesn't know your colleagues' names. Pass them in the initial prompt with --prompt "Marie, Julien, Samir, ACME Corp, ProjetX". They will be spelled correctly.
Compressed audio (Zoom, Teams)
Zoom MP4 recordings often contain poor-quality mono audio. Prefer Zoom's “audio only” export (M4A), or record in parallel with OBS if quality is a priority.
ggml_metal_init error on Mac
You compiled without Metal, or whisper.cpp cannot find the resources. Recompile from scratch: rm -rf build && cmake -B build && cmake --build build --config Release -j.
Ollama timeout on large context
The HTTP client cuts out before completion. Increase timeout=600 (10 min) on the requests side. On the server side, OLLAMA_KEEP_ALIVE=30m keeps the model warm.
Truncated transcription
If the summary seems to cut off the audio partway through, num_ctx is too low. Check the token count (wc -w on the .txt × 1.3 ≈ tokens) and set num_ctx to twice that.
Diarization (who is speaking)
whisper.cpp does not perform diarization. To distinguish speakers, add pyannote.audio to the pipeline before transcription, then annotate each segment. This is another layer, outside the scope here.
i
Actual hardware cost
For daily use—2 to 3 meetings transcribed per day—a Mac mini M4 16 GB (~€700) or a mini PC with RTX 3060 12 GB (~€600 used) is more than sufficient. It is a one-time investment versus ~€15/month/user for transcription SaaS, paying for itself in under a year for a team of 5.

#Go further

Your pipeline transcribes and summarizes locally. Three logical directions depending on your needs:

Medical report with faster-whisper
The consultation transcription guide uses faster-whisper (an optimized Python port) with a focus on medical confidentiality, pyannote diarization, and French healthcare GDPR compliance. An interesting variation on the same pattern.
Choose the right summarization LLM
Summarization is where quality varies the most. The guide to choosing your quantization compares Q4_K_M, Q5_K_M, and Q8_0—a noticeable difference on long, nuanced text.
Automate the pipeline with n8n
To trigger the chain as soon as an audio file arrives (watched folder, email, Slack), the n8n + Ollama automation guide shows you how to build the workflow without code.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.

Prices in euros (€) are French market prices including VAT, as checked by BestLLMfor. US prices differ: the Amazon buttons show the current US price.