Advanced 22 minHealth

Medical transcription and local LLM: data compliance patient

Dictate a consultation, get a clean transcript, then produce a SOAP note ready to paste into the patient's record—all without a single second of audio leaving the practice. That's the promise this guide makes operational: a local HDS-compliant medical transcription LLM pipeline built on Whisper-large-v3 and a 14B+ model (Llama 4 or Qwen3), designed to remain compliant with medical confidentiality and the HDS framework.

By Léa B.·Update 2026-06-15·Tested on Windows, macOS, and Linux

#HDS framework and medical confidentiality

Medical confidentiality (Article L.1110-4 of the French Public Health Code) applies to any document containing identifiable health information, including audio transcripts. As soon as patient data is processed outside the premises of the practice or hospital, the hosting provider must be HDS-certified (ASIP/ANS framework). In practice, sending a consultation to a non-HDS cloud API—OpenAI Whisper, Claude, or any mainstream public platform—puts you outside the legal framework, even for a test.

The answer: never let the audio or transcription leave the scope you control. That is precisely what a local LLM medical transcription pipeline enables: the inference machine is in the office, the disk is encrypted, and the network is isolated. No hosting provider, no authorization to request.

!
Non-HDS cloud = out of scope
Many practitioners use ChatGPT or hosted Whisper to save time. This violates professional confidentiality and may lead to disciplinary (Article 4 of the Code of Medical Ethics) and criminal (Article 226-13 of the French Penal Code) prosecution. An ARS audit can easily detect it through outbound traffic.

#100% local architecture

The AI at Work Kit

Deploy local AI at work: privacy, compliance, multi-user architecture, costs, the one-page memo for leadership.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

The pipeline consists of four components, all local: audio capture → Whisper-large-v3 → LLM with SOAP template → export to the practice-management software (LGC) or the DMP.

Capture
Lavalier mic or cardioid USB mic placed between practitioner and patient. No Bluetooth (latency and compression). Record in mono WAV at 16 kHz.
Transcription
Whisper-large-v3 (OpenAI, open weights, MIT license) run via faster-whisper on a local GPU. Model loaded into memory; audio destroyed after transcription.
Structuring
Llama 4 Scout (17B-A2B, MoE) or Qwen3-14B served by Ollama. The LLM consumes the transcription and produces a SOAP (Subjective, Objective, Assessment, Plan) note.
Export
The report is inserted into the EHR (Weda, Medistory, AxiSanté, Doctolib Pro, Maiia) via clipboard, an HL7 CDA shortcut, or an API if available.
i
Why Whisper-large-v3 instead of v3-turbo
Whisper-large-v3-turbo is faster but loses 1 to 2 WER points (Word Error Rate) on medical French, and especially degrades proper names and dosage instructions. In clinical settings, accuracy takes priority over speed—especially since the consultation has already ended when the pipeline is launched.

#Hardware and software requirements

GPU
RTX 4070 12 GB minimum, RTX 4080 16 GB or RTX 4090 24 GB recommended. Whisper-large-v3 uses ~3 GB of VRAM, while the 14B Q4_K_M LLM uses around 9 GB.
Mac alternative
A Mac mini M4 Pro with 48 GB of unified memory runs the full stack. Ideal for a professional practice: quiet, fanless when idle, and low-power.
Storage
LUKS-encrypted NVMe SSD (Linux), BitLocker (Windows Pro), or FileVault (macOS). 100 GB is enough: models ~25 GB, with the rest for temporary transcriptions.
Network
The inference workstation is on a VLAN separate from the patient Wi-Fi. No outbound Internet access in production—only during supervised updates.
Software
Ollama (≥ 0.5) to serve the LLM on http://localhost:11434, faster-whisper (Python) for transcription, and ffmpeg for audio conversion.
→
Progressive air gap
Start with a simple firewall rule blocking all outbound traffic except to update domains (ollama.com, huggingface.co, depots distrib). Disable the rule only during approved, logged maintenance windows.

#1. Whisper-large-v3 on French consultations

faster-whisper is a CTranslate2 fork that is 4 to 5× faster than the reference Python implementation, at identical quality. It accepts the same models and exposes a simple API.

Installation
python -m venv ~/venv-medtranscript
source ~/venv-medtranscript/bin/activate
pip install faster-whisper==1.1.0 ffmpeg-python

# Téléchargement du modèle (téléchargé une fois, ~3 Go)
python -c "from faster_whisper import WhisperModel; WhisperModel('large-v3', device='cuda', compute_type='float16')"

The transcription code itself fits in a few lines. Note the language='fr' setting, which disables automatic detection and prevents drift at the beginning of a consultation.

transcribe.py
from faster_whisper import WhisperModel
import sys, pathlib

AUDIO = pathlib.Path(sys.argv[1])
model = WhisperModel('large-v3', device='cuda', compute_type='float16')

segments, info = model.transcribe(
    str(AUDIO),
    language='fr',
    vad_filter=True,             # coupe les silences
    vad_parameters={'min_silence_duration_ms': 500},
    beam_size=5,
    initial_prompt=(
        'Consultation médicale en français. Vocabulaire clinique, '
        'noms de molécules, posologies. Termes courants : doliprane, '
        'amoxicilline, lévothyrox, ECG, IRM, NFS, CRP, HbA1c.'
    ),
)

transcript = '\n'.join(seg.text.strip() for seg in segments)
OUT = AUDIO.with_suffix('.txt')
OUT.write_text(transcript, encoding='utf-8')
print(f'Transcription : {OUT}')
→
The role of initial_prompt
Whisper conditions its decoding on this prompt. Listing domain vocabulary, common molecules, and biological acronyms significantly improves recognition—it's the low-tech version of fine-tuning. Adapt the list to your specialty (cardiology, pediatrics, gynecology).

On a RTX 4080, a 15-minute consultation is transcribed in about 90 seconds. The source audio file is deleted immediately afterward—the transcription is sufficient, and the audio creates unnecessary risk.

#2. LLM and SOAP template

The SOAP note is the de facto standard: Subjective (reason for visit, patient complaints), Objective (clinical examination, vital signs, additional tests), Assessment (diagnosis or hypotheses), Plan (prescriptions, upcoming tests, follow-up). The LLM is instructed to fill out this template strictly, without making anything up.

Model selection
# Option 1 — Llama 4 Scout (multimodal, 17B MoE, ~10 Go en Q4)
ollama pull llama4:scout

# Option 2 — Qwen3 14B (excellent en français)
ollama pull qwen3:14b

# Option 3 — Mistral Magistral 24B si VRAM ≥ 16 Go
ollama pull magistral:24b

The system prompt enforces a strict framework: no invention, literal citation when uncertain, and French medical language.

system_prompt_soap.txt
Tu es un assistant médical chargé de structurer une transcription
de consultation au format SOAP. Tu écris en français médical clair.

RÈGLES IMPÉRATIVES :
1. N'INVENTE JAMAIS de symptôme, diagnostic, traitement ou posologie
   qui ne figure pas explicitement dans la transcription.
2. En cas d'information ambiguë, écris littéralement :
   "À préciser par le praticien : [extrait textuel de la transcription]".
3. Conserve les unités exactes (mg, mL, fois/jour). Ne convertis rien.
4. Pour toute posologie, cite mot pour mot la phrase d'origine
   entre guillemets dans la section Plan.
5. Ne formule pas de diagnostic définitif si le praticien ne l'a pas
   posé. Utilise "hypothèse diagnostique" ou "à confirmer par...".
6. Garde uniquement ce qui relève de la consultation. Ignore les
   éléments hors propos (conversations annexes, interruptions).

STRUCTURE DE SORTIE (Markdown, exactement ces sections) :

## Subjective
- Motif de consultation :
- Histoire de la maladie actuelle :
- Antécédents pertinents évoqués :
- Traitements en cours évoqués :

## Objective
- Examen clinique :
- Constantes mentionnées :
- Examens complémentaires cités :

## Assessment
- Hypothèse(s) diagnostique(s) :
- Diagnostics différentiels évoqués :

## Plan
- Prescriptions (citer textuellement) :
- Examens complémentaires demandés :
- Conseils donnés au patient :
- Suivi prévu :

À la fin, ajoute une section :

## Points à vérifier par le praticien
Liste des éléments douteux, manquants ou nécessitant validation.
!
Dosage: zero tolerance for invention
An LLM-generated dosage error can kill. The rule requiring textual citations in quotation marks is non-negotiable. If the model rephrases a dose, you have a problem—change the model or make the prompt stricter.

#3. End-to-end pipeline

We chain transcription and then structuring through Ollama's local REST API. The script remains compact and auditable — about 40 lines, intentionally: less code means a smaller attack surface.

pipeline_soap.py
import sys, json, pathlib, requests, hashlib, datetime
from faster_whisper import WhisperModel

AUDIO = pathlib.Path(sys.argv[1])
MODEL_LLM = 'qwen3:14b'
SYS_PROMPT = pathlib.Path('system_prompt_soap.txt').read_text('utf-8')

# --- 1. Transcription locale
whisper = WhisperModel('large-v3', device='cuda', compute_type='float16')
segments, _ = whisper.transcribe(str(AUDIO), language='fr', vad_filter=True)
transcript = '\n'.join(s.text.strip() for s in segments)

# --- 2. Structuration SOAP via Ollama (localhost uniquement)
resp = requests.post(
    'http://localhost:11434/api/chat',
    json={
        'model': MODEL_LLM,
        'messages': [
            {'role': 'system', 'content': SYS_PROMPT},
            {'role': 'user', 'content': transcript},
        ],
        'options': {'temperature': 0.1, 'num_ctx': 8192},
        'stream': False,
    },
    timeout=600,
)
soap = resp.json()['message']['content']

# --- 3. Sortie + hash d'intégrité pour le journal d'audit
ts = datetime.datetime.now().isoformat(timespec='seconds')
sha = hashlib.sha256(soap.encode('utf-8')).hexdigest()[:16]
out = AUDIO.with_suffix('.soap.md')
out.write_text(
    f'<!-- généré le {ts} — modèle {MODEL_LLM} — sha256:{sha} -->\n\n{soap}',
    encoding='utf-8',
)

# --- 4. Effacement immédiat de l'audio et de la transcription brute
AUDIO.unlink()
print(f'Compte-rendu : {out}')
→
Low temperature, generous context
A temperature of 0.1 locks down creativity—which is desirable here. num_ctx 8192 covers a long consultation without loss. If you handle complex cases (psychiatry, neurology with an extensive medical history), increase it to 16384 and choose a model whose native context window supports it (Llama 4 Scout is comfortable up to 256k).

A 15-minute consultation completes the pipeline in under 3 minutes on RTX 4080. The .soap.md report is ready to paste into the LGC.

#4. Practice-management software integration

Three integration levels depending on your LGC:

  1. 01
    Level 1 — Clipboard
    The script automatically copies the report (pyperclip). The practitioner pastes it into the LGC in two clicks. Compatible with 100% of LGCs, with zero integration on the vendor's side. This is the recommended starting point.
  2. 02
    Level 2 — HL7 CDA shortcut
    The Markdown report is converted to CDA R2 (Clinical Document Architecture) by a pandoc script plus an XML template, then imported through the LGC's import documentaire function. Compatible with Weda, Medistory, and AxiSanté, which accept CDA.
  3. 03
    Level 3 — Native LGC API
    Some vendors (Doctolib Pro, Maiia) expose an observation-insertion API. The pipeline pushes the report directly into the patient record identified by INS. This is the most convenient option but requires a vendor agreement and CPS authentication.
i
National Health Identifier (INS)
If you automate linking to the correct patient, use the qualified INS, not the internal record number. It is the only key that survives a change of LGC and is legally enforceable. The INS is retrieved through the INSi teleservice.

#Compliance and audit logging

Use of a medical writing assistance tool must be documented. At least three items in the patient record:

AI-assisted writing disclosure
A line at the bottom of the report: "Structured report assisted by a local AI model (Qwen3-14B, 2026-04-12 version). Validated by Dr. [...] on [date]." This statement provides traceability.
Access log
Each pipeline run is logged locally (who, when, which audio, report hash). The log is append-only, signed, and retained for 10 years (the patient-record retention period).
GDPR DPIA
Conduct an impact assessment (Article 35 GDPR) for the “AI-assisted transcription” processing activity. Being 100% local simplifies the assessment: no transfer outside the EU, no processor, single purpose.
Patient information
The waiting-room poster and the notice in the intake form must indicate the use of a transcription-assistance tool. The patient may refuse—plan an alternative workflow.
Encryption at rest
The entire disk is encrypted. Reports are stored in the patient's LGC folder, not on the inference workstation's file system. Purge temporary .soap.md files at the end of the day.
Supervised updates
Every model update (Whisper, LLM) triggers revalidation on a set of 10 anonymized test queries. Known regression bug = no deployment.
!
Deleting a file ≠ secure erasure
unlink() removes the inode but leaves the blocks recoverable on an unencrypted SSD. Full-disk encryption (LUKS/FileVault/BitLocker) is what makes deletion irreversible in practice. Do not skip it.

#Go further

This pipeline covers the everyday needs of a private practitioner or a small hospital department. Three natural extensions: industrialize multi-workstation deployment with Docker and a CPS-authentication reverse proxy, add a local RAG over HAS recommendations to assist with Assessment, and harden network security all the way to a complete air gap.

Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.