Healthcare: transcription of consultations
Yes: Whisper (faster-whisper) transcribes the audio, and a local LLM drafts the report, without anything leaving the physician’s workstation. Local processing does not eliminate the applicable framework: professional confidentiality, informing the patient of their right to object, consent to voice recording, an HDS-certified host if a third party stores the data, and review by the physician themselves.
A doctor who writes up reports in the evening can save time with a local audio-to-transcription-to-structured-report pipeline. You’ll learn how to build it, understand its limitations (hallucinations, truncated context), and comply with the framework established by the HAS, CNIL, and the Ordre des médecins, while correcting several common misconceptions.
#Transcribing a consultation locally: what is permitted, and under what conditions
Yes, a doctor can transcribe consultations with AI without the audio leaving their workstation: Whisper transcribes the speech, a local LLM (via Ollama) drafts a structured report, and the professional reviews and approves it. This does not remove the legal obligations. The 2026 HAS and CNIL guide reminds us that introducing an AI system does not change the rules on professional confidentiality under Article L. 1110-4 of the French Public Health Code. You must also inform the patient and respect their right to object. Hosting by a third party requires an HDS-certified host. Local processing specifically avoids this transfer and its risks. This guide builds the pipeline, then details the framework and essential checks while correcting a few common misconceptions.
According to that same guide, automatic consultation transcription is the most common generative AI use case in healthcare. The issue, then, is not whether it can be done, but how to do it properly.
#The French framework: what the official guides say
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
The table covers each requirement and its concrete effect on a local pipeline. It also corrects two common misconceptions: patient consent is not the legal basis for processing, and HDS is not required by default, but when a third party hosts the data.
| Topic | What the guides say | Practical consequence |
|---|---|---|
| Trade secret | AI does not change the rules of Article L. 1110-4 of the CSP | The report remains confidential; limit access to the machine |
| Hosting | A third party that hosts health data (in the cloud) must be certified as a health data host | Processing on the physician’s workstation: no hosting provider in the workflow |
| Legal basis | CNIL considers that consent is neither the legal basis nor the exception for these processing activities | Document the selected legal basis with your DPO; don't rely on a checked box |
| Information | In most cases, fair, accessible information with the right to object is sufficient | Display or state that the consultation is being transcribed; respect the refusal |
| Voice recording | The use of identifying voice recordings is subject to consent under the French Civil Code | Obtain the patient's consent before recording their voice |
| Report validation | It should not be validated by a third party absent from the consultation, such as a secretary | The doctor who saw the patient rereads it personally |
| Retention | The Conseil de l'Ordre recommends 20 years after the last consultation, in the absence of a text specific to private practice | The report follows this timeline; there is no reason to keep the raw audio |
#Two misconceptions to abandon
First misconception: “ChatGPT is formally prohibited.” The rules do not target any particular tool. The issue is structural: sending health data to an online service moves it off your device, which calls for a certified hosting provider, a data-processing agreement, and appropriate notice. The HAS and CNIL guide also identifies a risk of violating medical confidentiality when sensitive data is integrated into prompts sent to online conversational agents.
Second misconception: “local means compliant.” Local processing eliminates the transfer, not the obligation to keep a register, inform patients, secure the workstation, and consider whether an impact assessment is required. The guide lists the DPIA among the deployer's obligations, where applicable: get support to decide this point, as well as whether your tool may qualify as a medical device.
#The stack: Whisper, a local LLM, and optional diarization
| Building block | Role | Good to know |
|---|---|---|
| faster-whisper | Transcribes audio (CTranslate2 implementation of Whisper) | Advertised as up to four times faster than the original implementation at equal precision |
| Large-v3 model | Multilingual model from the Whisper family, including French | Downloaded on first use; validate on your own recordings |
| Ollama + qwen3.5:9b | Write the structured report | 9-billion-parameter model; 6.6 GB according to the Ollama library |
| pyannote (optional) | Diarization: who speaks when | Model to download with a Hugging Face token after accepting its terms |
For throughput, the faster-whisper documentation provides a dated reference point: 13 minutes of audio transcribed in 1 minute 03 seconds with the large-v2 model in fp16, using 4,525 MB of VRAM on an RTX 3070 Ti with 8 GB (maintainer benchmark, CUDA 12.4). The same maintainers measured 59 seconds using 2,926 MB in int8. These are their measurements, on their machine, not ours: they indicate that a modest graphics card can transcribe faster than real time, without promising the same results on your system.
#Installation
With a NVIDIA GPU, faster-whisper requires the cuBLAS and cuDNN 9 libraries for CUDA 12. Without them, the program fails while loading the model: install them before testing, or switch to CPU mode.
#Transcribe audio
The script below sets the language, enables the silence filter (VAD), and timestamps each segment. Adapt the first line: device cuda with float16 for a NVIDIA GPU, device cpu with int8 for a Mac or a PC without a compatible card.
The hotwords parameter, exposed by faster-whisper’s transcription function, lets you suggest expected terms to the model: drug names, acronyms, and tests. It helps but does not guarantee accuracy: always review proper names and dosages.
#Whisper can write what nobody said
The large-v3 model card warns that predictions may contain text not present in the audio, which it calls hallucination. During a consultation, this can produce a plausible sentence invented during silence, a modified number, or a medication name that sounds similar but is wrong. The VAD filter reduces the risk during silence; it does not replace proofreading.
#Write the structured report
The second script sends the transcript to a local model through the Ollama API. Two settings matter: a low temperature to limit hallucinations, and a sufficient context window. Ollama’s documentation indicates a default window of approximately 4,000 tokens under 24 GiB of VRAM. A twenty-minute consultation represents around 3,000 words, or several thousand tokens (an estimate that varies with speaking rate): without an explicit num_ctx, the end of the transcript may be truncated. The example therefore sets 16,384 tokens.
This prompt applies the principles from the guide to system prompts: verifiable rules, an explicit output format, and defined behavior when information is missing. The instruction “Not mentioned” matters: it prevents the model from filling an empty section with a plausible assumption.
#Diarization: distinguishing doctor and patient
Diarization assigns each segment to a speaker (SPEAKER_00, SPEAKER_01). It is mainly useful when you record the consultation itself, with two speakers; for a dictation by the doctor alone, it is unnecessary. The pyannote repository currently recommends the community-1 pipeline: you must accept its terms of use on Hugging Face and create an access token to download it. Inference then runs locally.
#Typical law-firm workflow
- 01Inform the PatientBefore recording, state that the consultation or dictation will be transcribed by a local tool and that the patient may object. If you record their voice, obtain their consent. Document the information in the record.
- 02SaveTwo options: dictate 1 to 2 minutes of notes after the consultation (the patient is no longer being recorded), or record the conversation (verbal consent required, diarization useful). Dictation minimizes the data collected.
- 03Place the file on the target machineTransfer the audio (WAV, MP3, M4A) to a dedicated, encrypted folder. A small script monitors the folder and triggers processing.
- 04Run the pipelineThe script transcribes, summarizes, saves the report as text next to the audio, and logs the operation without writing any medical content to it.
- 05Review and validate it yourselfThe doctor who saw the patient reviews the report, corrects it, and then enters it into their practice-management software. A third party who was absent from the consultation cannot spot hallucinations.
- 06Delete the raw audioOnce the report is approved, delete the audio and intermediate transcript: the data no longer serves a purpose, and the report becomes part of the medical record.
#What to check during proofreading
| Item | Risk | Verification |
|---|---|---|
| Medications and dosages | Similar name, changed unit or dosage | Compare with the actual prescription |
| Background | An invented history or one attributed to the wrong patient | Compare against the dossier |
| Negations | “No fever” became “fever” | Review negative sentences |
| Empty sections | Plausible assumption substituted for “Not mentioned” | Check that nothing has been filled in |
| Dates and durations | Incorrect processing time or follow-up delay | Cross-check against what was said |
#Retention, security, and traceability
- Record of processing activities
- The CNIL and French Medical Council guide requires you to keep a record of processing activities; add a line describing this local pipeline and its purposes.
- Retention period
- The report follows the medical record's retention period (20 years after the last consultation, according to the professional body's recommendation). The raw audio, no longer needed once the report is validated, is deleted.
- Encryption and backups
- Encrypt the disk (FileVault, BitLocker, LUKS) and back it up regularly to encrypted media stored outside the office. A stolen workstation with unencrypted health data is a data breach.
- Traceability
- Keep a record of the patient's information, the tool versions, and the physician's validation: these are your evidence in the event of an audit.
- Medical transcription and local LLMs: patient data compliance
- Whisper + Ollama locally: 100% offline transcription
- Medical record summarization
- Encrypt the model drive
- WhisperX: word-level timestamps and speakers
- Master system prompts
- Source: HAS and CNIL guide on the proper use of AI systems in healthcare settings
- Source: CNOM and CNIL guide on protecting patient data
- Source: faster-whisper
- Source: the Whisper large-v3 model card
- Source: pyannote-audio
Can you transcribe a medical consultation with AI?+
Do you need the patient’s consent to transcribe their consultation?+
Is an HDS host required for a local pipeline?+
Is Whisper reliable for French medical vocabulary?+
Can a medical secretary approve the generated report?+
What configuration is needed to run this pipeline?+
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.