Intermediate 12 minDesktop

Automatic meeting minutes: Whisper + 100% LLM local

An AI meeting report has two building blocks: transcribe the audio, then summarize the text. The problem is that most SaaS tools send your entire meeting recording—names, numbers, strategy—to third-party servers. This guide builds the same pipeline entirely locally: Whisper for transcription, an Ollama LLM to extract decisions and action items, and nothing that leaves your machine.

By Marie L.·Update 2026-08-27·Tested on Windows, macOS, and Linux

#Why local is essential for sensitive meetings

A meeting is not an innocuous file. In one hour of recording, you may find client names, amounts, HR decisions, unannounced product directions, and sometimes personal data as defined by the GDPR. Entrusting that audio to an automated meeting-summary service in the cloud means accepting that all this content will be transferred, processed, and potentially retained by a third party.

Real privacy
Audio and transcription stay on your machine. No business data passes through an external server that could log it or use it for training.
Simplified GDPR compliance
No transfer to a subcontractor, often outside the EU. The “personal data transfer” section of your impact assessment disappears at the source.
Zero recurring cost
No per-user subscription or per-minute transcription billing. Once the machine is set up, you can process as many meetings as you want.
Works offline
A trip, a site without a reliable connection, a meeting in an isolated room: processing remains available without Internet.
i
Local doesn’t mean you can skip notifying people
Recording a meeting remains subject to rules: inform participants and obtain their consent. Local processing solves the data-transfer issue, not consent to recording.

#The two-stage pipeline principle

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

An AI-generated meeting summary always breaks down into two distinct steps that should not be confused. The first converts sound into text: that is Whisper's role, OpenAI's speech recognition model, for which several open-source implementations run perfectly locally. The second converts that raw text into a structured document: that is the role of a local LLM served by Ollama.

Transcription (Whisper)
Input: an audio or video file. Output: the meeting transcript in full, optionally with timestamps. It's computationally intensive but mechanical.
Summary (LLM Ollama)
Input: the transcript. Output: meeting notes—a summary, decisions made, and action items with their owner. This is where the prompt makes all the difference.

Separating the two steps has a practical advantage: you can rerun the synthesis as many times as you like with different prompts, without repeating the transcription, which is the slowest part.

#Prerequisites

Ollama installed
The daemon listens on http://localhost:11434 by default. If you haven't already, see the Ollama installation guide listed at the bottom of the page.
A summarization LLM
A 9B to 12B model in Q4_K_M (≈7 GB of VRAM) with good French performance is more than sufficient. Qwen 3.5 9B (256k context, perfect for long transcripts) or Gemma 4 12B are excellent choices for structured summarization.
Whisper
A local implementation: openai-whisper (simple), faster-whisper (fast), or whisper.cpp (lightweight, no GPU). This guide uses the first two.
ffmpeg
Essential for reading video formats (Teams/Zoom mp4 files) and converting audio. Whisper relies on it behind the scenes.
Hardware
A GPU greatly accelerates Whisper, but medium also runs on the CPU (more slowly). Allow ~2 GB of VRAM for the small model, and more for large-v3.
→
Which Whisper model to choose
For French, small already delivers good results on clean audio. Move to medium for a meeting with multiple speakers or average-quality audio, and to large-v3 when transcription quality matters more than speed (accents, jargon, overlapping speech).

#Step 1: transcribe the meeting audio with Whisper

The most direct installation uses the openai-whisper package, which provides a ready-to-use command-line command. Install it in a Python environment, with ffmpeg installed at the system level.

Installation
# ffmpeg (Debian/Ubuntu)
sudo apt install ffmpeg

# Whisper (dans un venv de préférence)
pip install -U openai-whisper

The transcription itself takes one command. Specify the model, language (forcing it prevents detection errors on short meetings), and text output format.

Terminal
whisper reunion.mp3 \
  --model medium \
  --language French \
  --output_format txt \
  --output_dir ./transcriptions

You get a reunion.txt file containing the entire meeting. If you also want timestamped subtitles (useful for finding a passage), add --output_format srt, or all to generate everything at once.

For a one-hour meeting, transcription can be slow on the CPU. That’s where faster-whisper comes in: a reimplementation that uses CTranslate2 and runs significantly faster at identical quality. Here’s a Python script that transcribes and saves the text.

transcrire.py
from faster_whisper import WhisperModel

# device="cuda" si GPU NVIDIA, sinon "cpu"
model = WhisperModel("medium", device="cuda", compute_type="float16")

segments, info = model.transcribe(
    "reunion.mp3",
    language="fr",
    vad_filter=True,   # coupe les silences, gagne du temps
)

with open("reunion.txt", "w", encoding="utf-8") as f:
    for seg in segments:
        f.write(seg.text.strip() + "\n")

print(f"Langue détectée : {info.language} — durée : {info.duration:.0f}s")
→
The vad_filter option
The VAD filter (voice activity detection) removes pauses and silence before transcription. In a real meeting, full of pauses and dead time, it significantly speeds up processing without degrading the useful text.

#Process a Teams or Zoom recording

Teams and Zoom recordings are .mp4 video files (sometimes .m4a for audio-only). Whisper can read them directly because it relies on ffmpeg, but extracting the audio track first as mono 16 kHz is more reliable and faster—especially with whisper.cpp, which requires it.

Extract audio from the .mp4
ffmpeg -i reunion_teams.mp4 \
  -ar 16000 -ac 1 \
  -c:a pcm_s16le \
  reunion.wav
-ar 16000
Resamples to 16 kHz, the frequency Whisper expects. There is no need to keep 48 kHz: speech does not require it.
-ac 1
Merge the channels into mono. There is no reason for a meeting recording to remain stereo for transcription.
-c:a pcm_s16le
Outputs uncompressed WAV, the safest format for Whisper input.

Where can you find the file? In Teams, recordings land in OneDrive/SharePoint (the "Recordings" folder): download the .mp4 locally before processing. In Zoom, the local recording is in the Documents/Zoom folder, with a separate audio.m4a file if the option is enabled—that's the one to target directly, without going through the video.

!
Do not process from the SaaS cloud
The whole point of running locally is lost if you let the provider’s AI (Teams Premium, Zoom AI Companion) generate the minutes on the server side. Retrieve the raw file and process it on your machine: that is the only way to guarantee that nothing leaks.

#Step 2: synthesis by a local LLM

The raw transcript is unreadable: no reliable punctuation, no structure, and repetitions. The local LLM turns it into a usable report. The text is sent to Ollama with precise instructions about the expected output format.

The simplest way to test it: pipe the transcript into ollama run. The model receives the prompt followed by the meeting text.

Quick summary
ollama run qwen3.5:9b "Résume ce compte rendu de réunion en français, \
en listant les points clés, les décisions et les actions à mener. \
Voici la transcription :

$(cat reunion.txt)"

For recurring use, it's better to use Ollama's REST API on port 11434: it cleanly separates the system message (the role and format) from the content (the transcription), making the result reusable in a script.

synthese.py
import requests

transcript = open("reunion.txt", encoding="utf-8").read()

SYSTEM = """Tu es un assistant qui rédige des comptes rendus de réunion clairs et concis en français.
Tu ne inventes rien : tu ne t'appuies que sur la transcription fournie."""

resp = requests.post("http://localhost:11434/api/chat", json={
    "model": "qwen3.5:9b",
    "stream": False,
    "messages": [
        {"role": "system", "content": SYSTEM},
        {"role": "user", "content": PROMPT + "\n\n---\n" + transcript},
    ],
})

print(resp.json()["message"]["content"])
i
Watch out for the context window
A one-hour meeting can easily produce 8,000 to 12,000 words. Make sure the model's context is large enough (num_ctx). Beyond that, split the transcript into sections and summarize them in chunks before a final synthesis—a mini map-reduce pipeline.

#The prompt that cleanly extracts decisions and actions

This is the core of a good AI meeting summary. A vague prompt produces a wall of summary text; a structured prompt produces a document you can use directly. The key is to require an explicit output format, organized into sections, and ask the model to clearly distinguish what has been decided from what remains to be done.

Synthesis prompt
À partir de la transcription de réunion ci-dessous, produis un compte rendu structuré en français, en respectant EXACTEMENT ce format :

## Résumé
Trois à cinq phrases sur l'objet et les conclusions de la réunion.

## Décisions prises
- Une puce par décision actée, formulée clairement.

## Actions à mener
- [Responsable] Action précise — échéance si mentionnée.

## Points en suspens
- Sujets évoqués sans conclusion, à traiter plus tard.

Règles :
- Ne reprends que ce qui est réellement dit dans la transcription.
- Si le responsable ou l'échéance d'une action n'est pas mentionné, écris « non précisé ».
- Reste factuel, pas de reformulation marketing.
A required format
Section headings (##) force the model to sort the information instead of mixing everything together. You get consistent output from one meeting to the next.
The anti-hallucination rule
“Only repeat what is actually stated” and “not specified” greatly reduce the risk that the LLM will invent a deadline or an owner who doesn't exist.
The person responsible in brackets
Putting [Owner] at the start of an action makes the meeting notes immediately actionable—you can see who does what at a glance.
→
Test several models on the same transcript
Since the transcription is already done, comparing models costs nothing: replay the summary with Qwen 3.5 9B and then Gemma 4 12B on the same reunion.txt. You'll quickly see which one structures actions better in French.

#Automate the entire pipeline

Once both steps have been validated, chain them together in a single script: give it a meeting file, and it produces the minutes in Markdown. That's what turns the process into a daily tool.

compte-rendu.sh
#!/usr/bin/env bash
set -euo pipefail

INPUT="$1"                 # reunion.mp4 ou reunion.mp3
BASE="$(basename "${INPUT%.*}")"

# 1. Extraire l'audio en 16 kHz mono
ffmpeg -y -i "$INPUT" -ar 16000 -ac 1 -c:a pcm_s16le "$BASE.wav"

# 2. Transcrire
whisper "$BASE.wav" --model medium --language French \
  --output_format txt --output_dir .

# 3. Synthétiser avec le LLM local
python synthese.py "$BASE.txt" > "$BASE-compte-rendu.md"

echo "Compte rendu prêt : $BASE-compte-rendu.md"

Adapt synthese.py so it reads the file passed as an argument and writes the result to standard output. You then have a single command: ./compte-rendu.sh reunion_teams.mp4, and the Markdown report appears a few minutes later, without a single byte leaving the machine.

#Troubleshooting

Transcription switches to English
Whisper misdetected the language because the beginning was silent or bilingual. Always force --language French (or language="fr").
Misspelled proper names
Normal: Whisper guesses phonetically. Add a quick correction step to the synthesis prompt ("correct names that are clearly mistranscribed") or provide the LLM with a glossary of names.
Meeting too long for the context
Split the transcript into blocks of ~3,000 words, summarize each one, then synthesize the summaries. Also increase num_ctx in the Ollama call if your VRAM allows it.
Whisper is very slow
Switch to faster-whisper, enable vad_filter, reduce the model (medium → small), or switch to GPU. whisper.cpp is also a good option on a machine without a GPU.
Overlapping voices
Whisper does not separate speakers. For a transcript showing “who said what,” you need to add a diarization step beforehand (e.g., pyannote)—a topic in its own right.

#Go further

This guide combines two building blocks already covered in detail on the site. To explore each step further or secure the whole setup:

Whisper + Ollama: 100% local transcription and summarization pipeline
The reference guide to the audio component, with Whisper implementation variants and an end-to-end one-hour meeting example.
Choose your quantization (Q4, Q5, Q8, FP16)
To fit a 14B or 32B summarization model on your GPU without degrading the quality of the summary.
Local LLMs and GDPR: private-data compliance for businesses
The regulatory angle: processing meetings locally simplifies compliance; this guide details what still needs to be documented.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.