Intermediate 11 minHealth

Healthcare: transcription of consultations

Direct response

Yes: Whisper (faster-whisper) transcribes the audio, and a local LLM drafts the report, without anything leaving the physician’s workstation. Local processing does not eliminate the applicable framework: professional confidentiality, informing the patient of their right to object, consent to voice recording, an HDS-certified host if a third party stores the data, and review by the physician themselves.

A doctor who writes up reports in the evening can save time with a local audio-to-transcription-to-structured-report pipeline. You’ll learn how to build it, understand its limitations (hallucinations, truncated context), and comply with the framework established by the HAS, CNIL, and the Ordre des médecins, while correcting several common misconceptions.

By Mohamed Meguedmi·Update 2026-09-30·Tested on Windows, macOS, and Linux

#Transcribing a consultation locally: what is permitted, and under what conditions

Yes, a doctor can transcribe consultations with AI without the audio leaving their workstation: Whisper transcribes the speech, a local LLM (via Ollama) drafts a structured report, and the professional reviews and approves it. This does not remove the legal obligations. The 2026 HAS and CNIL guide reminds us that introducing an AI system does not change the rules on professional confidentiality under Article L. 1110-4 of the French Public Health Code. You must also inform the patient and respect their right to object. Hosting by a third party requires an HDS-certified host. Local processing specifically avoids this transfer and its risks. This guide builds the pipeline, then details the framework and essential checks while correcting a few common misconceptions.

According to that same guide, automatic consultation transcription is the most common generative AI use case in healthcare. The issue, then, is not whether it can be done, but how to do it properly.

!
This guide is not legal advice
The legal points are those covered by the guides published by HAS, CNIL, and the Conseil national de l'Ordre des médecins. For your situation (solo practice, group practice, facility), have your DPO or professional regulatory body validate the setup.

#The French framework: what the official guides say

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

The table covers each requirement and its concrete effect on a local pipeline. It also corrects two common misconceptions: patient consent is not the legal basis for processing, and HDS is not required by default, but when a third party hosts the data.

Requirements and consequences for a local pipeline
TopicWhat the guides sayPractical consequence
Trade secretAI does not change the rules of Article L. 1110-4 of the CSPThe report remains confidential; limit access to the machine
HostingA third party that hosts health data (in the cloud) must be certified as a health data hostProcessing on the physician’s workstation: no hosting provider in the workflow
Legal basisCNIL considers that consent is neither the legal basis nor the exception for these processing activitiesDocument the selected legal basis with your DPO; don't rely on a checked box
InformationIn most cases, fair, accessible information with the right to object is sufficientDisplay or state that the consultation is being transcribed; respect the refusal
Voice recordingThe use of identifying voice recordings is subject to consent under the French Civil CodeObtain the patient's consent before recording their voice
Report validationIt should not be validated by a third party absent from the consultation, such as a secretaryThe doctor who saw the patient rereads it personally
RetentionThe Conseil de l'Ordre recommends 20 years after the last consultation, in the absence of a text specific to private practiceThe report follows this timeline; there is no reason to keep the raw audio

#Two misconceptions to abandon

First misconception: “ChatGPT is formally prohibited.” The rules do not target any particular tool. The issue is structural: sending health data to an online service moves it off your device, which calls for a certified hosting provider, a data-processing agreement, and appropriate notice. The HAS and CNIL guide also identifies a risk of violating medical confidentiality when sensitive data is integrated into prompts sent to online conversational agents.

Second misconception: “local means compliant.” Local processing eliminates the transfer, not the obligation to keep a register, inform patients, secure the workstation, and consider whether an impact assessment is required. The guide lists the DPIA among the deployer's obligations, where applicable: get support to decide this point, as well as whether your tool may qualify as a medical device.

#The stack: Whisper, a local LLM, and optional diarization

The pipeline building blocks
Building blockRoleGood to know
faster-whisperTranscribes audio (CTranslate2 implementation of Whisper)Advertised as up to four times faster than the original implementation at equal precision
Large-v3 modelMultilingual model from the Whisper family, including FrenchDownloaded on first use; validate on your own recordings
Ollama + qwen3.5:9bWrite the structured report9-billion-parameter model; 6.6 GB according to the Ollama library
pyannote (optional)Diarization: who speaks whenModel to download with a Hugging Face token after accepting its terms

For throughput, the faster-whisper documentation provides a dated reference point: 13 minutes of audio transcribed in 1 minute 03 seconds with the large-v2 model in fp16, using 4,525 MB of VRAM on an RTX 3070 Ti with 8 GB (maintainer benchmark, CUDA 12.4). The same maintainers measured 59 seconds using 2,926 MB in int8. These are their measurements, on their machine, not ours: they indicate that a modest graphics card can transcribe faster than real time, without promising the same results on your system.

i
No CUDA on Mac
The official faster-whisper example shows execution on a processor in int8. This is the configuration to use on a Mac: slower than a NVIDIA GPU, but sufficient for deferred transcription after the consultation.

#Installation

Terminal
python3 -m venv venv
source venv/bin/activate

pip install faster-whisper ollama
# Optionnel pour la diarisation :
pip install pyannote.audio
Terminal
# Le modèle Whisper se télécharge au premier usage.
# Modèle de rédaction via Ollama :
ollama pull qwen3.5:9b

With a NVIDIA GPU, faster-whisper requires the cuBLAS and cuDNN 9 libraries for CUDA 12. Without them, the program fails while loading the model: install them before testing, or switch to CPU mode.

#Transcribe audio

The script below sets the language, enables the silence filter (VAD), and timestamps each segment. Adapt the first line: device cuda with float16 for a NVIDIA GPU, device cpu with int8 for a Mac or a PC without a compatible card.

Python
from faster_whisper import WhisperModel

# GPU NVIDIA : device="cuda", compute_type="float16"
# Sans GPU compatible : device="cpu", compute_type="int8"
model = WhisperModel("large-v3", device="cuda", compute_type="float16")

MOTS_CLES = "amoxicilline, paracétamol, HbA1c, tension artérielle"

def transcribe(audio_path):
    segments, info = model.transcribe(
        audio_path,
        language="fr",
        beam_size=5,
        vad_filter=True,        # coupe les silences
        hotwords=MOTS_CLES,     # vocabulaire attendu
        word_timestamps=False,
    )
    lines = []
    for s in segments:
        ts = f"[{int(s.start // 60):02d}:{int(s.start % 60):02d}]"
        lines.append(f"{ts} {s.text.strip()}")
    return "\n".join(lines)

if __name__ == "__main__":
    import sys
    print(transcribe(sys.argv[1]))

The hotwords parameter, exposed by faster-whisper’s transcription function, lets you suggest expected terms to the model: drug names, acronyms, and tests. It helps but does not guarantee accuracy: always review proper names and dosages.

#Whisper can write what nobody said

The large-v3 model card warns that predictions may contain text not present in the audio, which it calls hallucination. During a consultation, this can produce a plausible sentence invented during silence, a modified number, or a medication name that sounds similar but is wrong. The VAD filter reduces the risk during silence; it does not replace proofreading.

#Write the structured report

The second script sends the transcript to a local model through the Ollama API. Two settings matter: a low temperature to limit hallucinations, and a sufficient context window. Ollama’s documentation indicates a default window of approximately 4,000 tokens under 24 GiB of VRAM. A twenty-minute consultation represents around 3,000 words, or several thousand tokens (an estimate that varies with speaking rate): without an explicit num_ctx, the end of the transcript may be truncated. The example therefore sets 16,384 tokens.

Python
import requests

SYSTEM_CR = """Tu es un assistant pour médecin généraliste français.
À partir d'une TRANSCRIPTION BRUTE d'une consultation, tu produis un
compte-rendu médical structuré.

FORMAT DE SORTIE :
## Motif de consultation
## Antécédents (mentionnés)
## Examen clinique
## Diagnostic / Hypothèses
## Traitement / Prescriptions
## Suivi

RÈGLES :
- Reste TEXTUEL : n'ajoute aucun élément non mentionné dans la transcription.
- Reformule en vocabulaire médical standard.
- Si une section n'est pas abordée, écris "Non mentionné".
- Ne complète jamais une posologie ou un diagnostic absent.
"""

def summarize(transcription):
    r = requests.post("http://localhost:11434/api/chat", json={
        "model": "qwen3.5:9b",
        "messages": [
            {"role": "system", "content": SYSTEM_CR},
            {"role": "user",   "content": transcription},
        ],
        "stream": False,
        "options": {"temperature": 0.2, "num_ctx": 16384},
    })
    return r.json()["message"]["content"]

This prompt applies the principles from the guide to system prompts: verifiable rules, an explicit output format, and defined behavior when information is missing. The instruction “Not mentioned” matters: it prevents the model from filling an empty section with a plausible assumption.

#Diarization: distinguishing doctor and patient

Diarization assigns each segment to a speaker (SPEAKER_00, SPEAKER_01). It is mainly useful when you record the consultation itself, with two speakers; for a dictation by the doctor alone, it is unnecessary. The pyannote repository currently recommends the community-1 pipeline: you must accept its terms of use on Hugging Face and create an access token to download it. Inference then runs locally.

Python
import torch
from pyannote.audio import Pipeline

pipeline = Pipeline.from_pretrained(
    "pyannote/speaker-diarization-community-1",
    token="HUGGINGFACE_ACCESS_TOKEN",
)
pipeline.to(torch.device("cuda"))  # si un GPU est disponible

output = pipeline("consultation.wav")
for turn, speaker in output.speaker_diarization:
    print(f"{turn.start:.1f}s - {turn.end:.1f}s : {speaker}")
→
Merge the two results
For a two-speaker transcript, align Whisper and pyannote segments by their timestamps before sending the text to the LLM. The WhisperX guide describes a tool that performs this alignment word by word.

#Typical law-firm workflow

  1. 01
    Inform the Patient
    Before recording, state that the consultation or dictation will be transcribed by a local tool and that the patient may object. If you record their voice, obtain their consent. Document the information in the record.
  2. 02
    Save
    Two options: dictate 1 to 2 minutes of notes after the consultation (the patient is no longer being recorded), or record the conversation (verbal consent required, diarization useful). Dictation minimizes the data collected.
  3. 03
    Place the file on the target machine
    Transfer the audio (WAV, MP3, M4A) to a dedicated, encrypted folder. A small script monitors the folder and triggers processing.
  4. 04
    Run the pipeline
    The script transcribes, summarizes, saves the report as text next to the audio, and logs the operation without writing any medical content to it.
  5. 05
    Review and validate it yourself
    The doctor who saw the patient reviews the report, corrects it, and then enters it into their practice-management software. A third party who was absent from the consultation cannot spot hallucinations.
  6. 06
    Delete the raw audio
    Once the report is approved, delete the audio and intermediate transcript: the data no longer serves a purpose, and the report becomes part of the medical record.

#What to check during proofreading

Review checklist for a generated report
ItemRiskVerification
Medications and dosagesSimilar name, changed unit or dosageCompare with the actual prescription
BackgroundAn invented history or one attributed to the wrong patientCompare against the dossier
Negations“No fever” became “fever”Review negative sentences
Empty sectionsPlausible assumption substituted for “Not mentioned”Check that nothing has been filled in
Dates and durationsIncorrect processing time or follow-up delayCross-check against what was said
Record of processing activities
The CNIL and French Medical Council guide requires you to keep a record of processing activities; add a line describing this local pipeline and its purposes.
Retention period
The report follows the medical record's retention period (20 years after the last consultation, according to the professional body's recommendation). The raw audio, no longer needed once the report is validated, is deleted.
Encryption and backups
Encrypt the disk (FileVault, BitLocker, LUKS) and back it up regularly to encrypted media stored outside the office. A stolen workstation with unencrypted health data is a data breach.
Traceability
Keep a record of the patient's information, the tool versions, and the physician's validation: these are your evidence in the event of an audit.
FAQ
Can you transcribe a medical consultation with AI?+
Yes, provided you follow the framework: professional confidentiality still applies, the patient must be informed and able to object, recording their voice requires their consent, and the report must be reviewed by the doctor. Local processing avoids transferring data to a third-party host, but does not eliminate the other obligations.
Do you need the patient’s consent to transcribe their consultation?+
For data processing, the CNIL considers consent to be neither the legal basis nor the exception: information and the right to object generally suffice. However, recording an identifiable person’s voice requires their consent under the Civil Code. Dictating your notes after the consultation avoids the issue.
Is an HDS host required for a local pipeline?+
The requirement applies to health data entrusted to a third party that hosts it, for example in the cloud. Processing performed on the doctor's workstation, without transfer, does not put a hosting third party in the flow. If you add a remote backup or online service, the question arises again.
Is Whisper reliable for French medical vocabulary?+
Whisper large-v3 is the reference model in the family, but its model card notes that it can produce text not present in the audio. Medication names, acronyms, and dosages are the weak points. The hotwords parameter helps, proofreading remains essential, and no one should approve it in place of a doctor.
Can a medical secretary approve the generated report?+
The HAS and CNIL guide recommends that the report not be validated by a third party who was absent from the consultation, because they will not be able to detect hallucinations. The physician who examined the patient reviews and validates it. The secretary can then file or integrate the document.
What configuration is needed to run this pipeline?+
A NVIDIA GPU with 8 GB of VRAM transcribes faster than real time according to faster-whisper measurements, and a 9-billion-parameter writing model requires 6.6 GB according to the Ollama library. Without a GPU, int8 CPU mode works, but more slowly: plan for batch processing.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.