Automatic meeting minutes: Whisper + 100% LLM local
An AI meeting report has two building blocks: transcribe the audio, then summarize the text. The problem is that most SaaS tools send your entire meeting recording—names, numbers, strategy—to third-party servers. This guide builds the same pipeline entirely locally: Whisper for transcription, an Ollama LLM to extract decisions and action items, and nothing that leaves your machine.
#Why local is essential for sensitive meetings
A meeting is not an innocuous file. In one hour of recording, you may find client names, amounts, HR decisions, unannounced product directions, and sometimes personal data as defined by the GDPR. Entrusting that audio to an automated meeting-summary service in the cloud means accepting that all this content will be transferred, processed, and potentially retained by a third party.
- Real privacy
- Audio and transcription stay on your machine. No business data passes through an external server that could log it or use it for training.
- Simplified GDPR compliance
- No transfer to a subcontractor, often outside the EU. The “personal data transfer” section of your impact assessment disappears at the source.
- Zero recurring cost
- No per-user subscription or per-minute transcription billing. Once the machine is set up, you can process as many meetings as you want.
- Works offline
- A trip, a site without a reliable connection, a meeting in an isolated room: processing remains available without Internet.
#The two-stage pipeline principle
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
An AI-generated meeting summary always breaks down into two distinct steps that should not be confused. The first converts sound into text: that is Whisper's role, OpenAI's speech recognition model, for which several open-source implementations run perfectly locally. The second converts that raw text into a structured document: that is the role of a local LLM served by Ollama.
- Transcription (Whisper)
- Input: an audio or video file. Output: the meeting transcript in full, optionally with timestamps. It's computationally intensive but mechanical.
- Summary (LLM Ollama)
- Input: the transcript. Output: meeting notes—a summary, decisions made, and action items with their owner. This is where the prompt makes all the difference.
Separating the two steps has a practical advantage: you can rerun the synthesis as many times as you like with different prompts, without repeating the transcription, which is the slowest part.
#Prerequisites
- Ollama installed
- The daemon listens on http://localhost:11434 by default. If you haven't already, see the Ollama installation guide listed at the bottom of the page.
- A summarization LLM
- A 9B to 12B model in Q4_K_M (≈7 GB of VRAM) with good French performance is more than sufficient. Qwen 3.5 9B (256k context, perfect for long transcripts) or Gemma 4 12B are excellent choices for structured summarization.
- Whisper
- A local implementation: openai-whisper (simple), faster-whisper (fast), or whisper.cpp (lightweight, no GPU). This guide uses the first two.
- ffmpeg
- Essential for reading video formats (Teams/Zoom mp4 files) and converting audio. Whisper relies on it behind the scenes.
- Hardware
- A GPU greatly accelerates Whisper, but medium also runs on the CPU (more slowly). Allow ~2 GB of VRAM for the small model, and more for large-v3.
#Step 1: transcribe the meeting audio with Whisper
The most direct installation uses the openai-whisper package, which provides a ready-to-use command-line command. Install it in a Python environment, with ffmpeg installed at the system level.
The transcription itself takes one command. Specify the model, language (forcing it prevents detection errors on short meetings), and text output format.
You get a reunion.txt file containing the entire meeting. If you also want timestamped subtitles (useful for finding a passage), add --output_format srt, or all to generate everything at once.
For a one-hour meeting, transcription can be slow on the CPU. That’s where faster-whisper comes in: a reimplementation that uses CTranslate2 and runs significantly faster at identical quality. Here’s a Python script that transcribes and saves the text.
#Process a Teams or Zoom recording
Teams and Zoom recordings are .mp4 video files (sometimes .m4a for audio-only). Whisper can read them directly because it relies on ffmpeg, but extracting the audio track first as mono 16 kHz is more reliable and faster—especially with whisper.cpp, which requires it.
- -ar 16000
- Resamples to 16 kHz, the frequency Whisper expects. There is no need to keep 48 kHz: speech does not require it.
- -ac 1
- Merge the channels into mono. There is no reason for a meeting recording to remain stereo for transcription.
- -c:a pcm_s16le
- Outputs uncompressed WAV, the safest format for Whisper input.
Where can you find the file? In Teams, recordings land in OneDrive/SharePoint (the "Recordings" folder): download the .mp4 locally before processing. In Zoom, the local recording is in the Documents/Zoom folder, with a separate audio.m4a file if the option is enabled—that's the one to target directly, without going through the video.
#Step 2: synthesis by a local LLM
The raw transcript is unreadable: no reliable punctuation, no structure, and repetitions. The local LLM turns it into a usable report. The text is sent to Ollama with precise instructions about the expected output format.
The simplest way to test it: pipe the transcript into ollama run. The model receives the prompt followed by the meeting text.
For recurring use, it's better to use Ollama's REST API on port 11434: it cleanly separates the system message (the role and format) from the content (the transcription), making the result reusable in a script.
#The prompt that cleanly extracts decisions and actions
This is the core of a good AI meeting summary. A vague prompt produces a wall of summary text; a structured prompt produces a document you can use directly. The key is to require an explicit output format, organized into sections, and ask the model to clearly distinguish what has been decided from what remains to be done.
- A required format
- Section headings (##) force the model to sort the information instead of mixing everything together. You get consistent output from one meeting to the next.
- The anti-hallucination rule
- “Only repeat what is actually stated” and “not specified” greatly reduce the risk that the LLM will invent a deadline or an owner who doesn't exist.
- The person responsible in brackets
- Putting [Owner] at the start of an action makes the meeting notes immediately actionable—you can see who does what at a glance.
#Automate the entire pipeline
Once both steps have been validated, chain them together in a single script: give it a meeting file, and it produces the minutes in Markdown. That's what turns the process into a daily tool.
Adapt synthese.py so it reads the file passed as an argument and writes the result to standard output. You then have a single command: ./compte-rendu.sh reunion_teams.mp4, and the Markdown report appears a few minutes later, without a single byte leaving the machine.
#Troubleshooting
- Transcription switches to English
- Whisper misdetected the language because the beginning was silent or bilingual. Always force --language French (or language="fr").
- Misspelled proper names
- Normal: Whisper guesses phonetically. Add a quick correction step to the synthesis prompt ("correct names that are clearly mistranscribed") or provide the LLM with a glossary of names.
- Meeting too long for the context
- Split the transcript into blocks of ~3,000 words, summarize each one, then synthesize the summaries. Also increase num_ctx in the Ollama call if your VRAM allows it.
- Whisper is very slow
- Switch to faster-whisper, enable vad_filter, reduce the model (medium → small), or switch to GPU. whisper.cpp is also a good option on a machine without a GPU.
- Overlapping voices
- Whisper does not separate speakers. For a transcript showing “who said what,” you need to add a diarization step beforehand (e.g., pyannote)—a topic in its own right.
#Go further
This guide combines two building blocks already covered in detail on the site. To explore each step further or secure the whole setup:
- Whisper + Ollama: 100% local transcription and summarization pipeline
- The reference guide to the audio component, with Whisper implementation variants and an end-to-end one-hour meeting example.
- Choose your quantization (Q4, Q5, Q8, FP16)
- To fit a 14B or 32B summarization model on your GPU without degrading the quality of the summary.
- Local LLMs and GDPR: private-data compliance for businesses
- The regulatory angle: processing meetings locally simplifies compliance; this guide details what still needs to be documented.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.