Medical transcription and local LLM: data compliance patient
Dictate a consultation, get a clean transcript, then produce a SOAP note ready to paste into the patient's record—all without a single second of audio leaving the practice. That's the promise this guide makes operational: a local HDS-compliant medical transcription LLM pipeline built on Whisper-large-v3 and a 14B+ model (Llama 4 or Qwen3), designed to remain compliant with medical confidentiality and the HDS framework.
#HDS framework and medical confidentiality
Medical confidentiality (Article L.1110-4 of the French Public Health Code) applies to any document containing identifiable health information, including audio transcripts. As soon as patient data is processed outside the premises of the practice or hospital, the hosting provider must be HDS-certified (ASIP/ANS framework). In practice, sending a consultation to a non-HDS cloud API—OpenAI Whisper, Claude, or any mainstream public platform—puts you outside the legal framework, even for a test.
The answer: never let the audio or transcription leave the scope you control. That is precisely what a local LLM medical transcription pipeline enables: the inference machine is in the office, the disk is encrypted, and the network is isolated. No hosting provider, no authorization to request.
#100% local architecture
Deploy local AI at work: privacy, compliance, multi-user architecture, costs, the one-page memo for leadership.
- Lifetime online access
- PDF + files
- Lifetime updates
The pipeline consists of four components, all local: audio capture → Whisper-large-v3 → LLM with SOAP template → export to the practice-management software (LGC) or the DMP.
- Capture
- Lavalier mic or cardioid USB mic placed between practitioner and patient. No Bluetooth (latency and compression). Record in mono WAV at 16 kHz.
- Transcription
- Whisper-large-v3 (OpenAI, open weights, MIT license) run via faster-whisper on a local GPU. Model loaded into memory; audio destroyed after transcription.
- Structuring
- Llama 4 Scout (17B-A2B, MoE) or Qwen3-14B served by Ollama. The LLM consumes the transcription and produces a SOAP (Subjective, Objective, Assessment, Plan) note.
- Export
- The report is inserted into the EHR (Weda, Medistory, AxiSanté, Doctolib Pro, Maiia) via clipboard, an HL7 CDA shortcut, or an API if available.
#Hardware and software requirements
- GPU
- RTX 4070 12 GB minimum, RTX 4080 16 GB or RTX 4090 24 GB recommended. Whisper-large-v3 uses ~3 GB of VRAM, while the 14B Q4_K_M LLM uses around 9 GB.
- Mac alternative
- A Mac mini M4 Pro with 48 GB of unified memory runs the full stack. Ideal for a professional practice: quiet, fanless when idle, and low-power.
- Storage
- LUKS-encrypted NVMe SSD (Linux), BitLocker (Windows Pro), or FileVault (macOS). 100 GB is enough: models ~25 GB, with the rest for temporary transcriptions.
- Network
- The inference workstation is on a VLAN separate from the patient Wi-Fi. No outbound Internet access in production—only during supervised updates.
- Software
- Ollama (≥ 0.5) to serve the LLM on http://localhost:11434, faster-whisper (Python) for transcription, and ffmpeg for audio conversion.
#1. Whisper-large-v3 on French consultations
faster-whisper is a CTranslate2 fork that is 4 to 5× faster than the reference Python implementation, at identical quality. It accepts the same models and exposes a simple API.
The transcription code itself fits in a few lines. Note the language='fr' setting, which disables automatic detection and prevents drift at the beginning of a consultation.
On a RTX 4080, a 15-minute consultation is transcribed in about 90 seconds. The source audio file is deleted immediately afterward—the transcription is sufficient, and the audio creates unnecessary risk.
#2. LLM and SOAP template
The SOAP note is the de facto standard: Subjective (reason for visit, patient complaints), Objective (clinical examination, vital signs, additional tests), Assessment (diagnosis or hypotheses), Plan (prescriptions, upcoming tests, follow-up). The LLM is instructed to fill out this template strictly, without making anything up.
The system prompt enforces a strict framework: no invention, literal citation when uncertain, and French medical language.
#3. End-to-end pipeline
We chain transcription and then structuring through Ollama's local REST API. The script remains compact and auditable — about 40 lines, intentionally: less code means a smaller attack surface.
A 15-minute consultation completes the pipeline in under 3 minutes on RTX 4080. The .soap.md report is ready to paste into the LGC.
#4. Practice-management software integration
Three integration levels depending on your LGC:
- 01Level 1 — ClipboardThe script automatically copies the report (pyperclip). The practitioner pastes it into the LGC in two clicks. Compatible with 100% of LGCs, with zero integration on the vendor's side. This is the recommended starting point.
- 02Level 2 — HL7 CDA shortcutThe Markdown report is converted to CDA R2 (Clinical Document Architecture) by a pandoc script plus an XML template, then imported through the LGC's import documentaire function. Compatible with Weda, Medistory, and AxiSanté, which accept CDA.
- 03Level 3 — Native LGC APISome vendors (Doctolib Pro, Maiia) expose an observation-insertion API. The pipeline pushes the report directly into the patient record identified by INS. This is the most convenient option but requires a vendor agreement and CPS authentication.
#Compliance and audit logging
Use of a medical writing assistance tool must be documented. At least three items in the patient record:
- AI-assisted writing disclosure
- A line at the bottom of the report: "Structured report assisted by a local AI model (Qwen3-14B, 2026-04-12 version). Validated by Dr. [...] on [date]." This statement provides traceability.
- Access log
- Each pipeline run is logged locally (who, when, which audio, report hash). The log is append-only, signed, and retained for 10 years (the patient-record retention period).
- GDPR DPIA
- Conduct an impact assessment (Article 35 GDPR) for the “AI-assisted transcription” processing activity. Being 100% local simplifies the assessment: no transfer outside the EU, no processor, single purpose.
- Patient information
- The waiting-room poster and the notice in the intake form must indicate the use of a transcription-assistance tool. The patient may refuse—plan an alternative workflow.
- Encryption at rest
- The entire disk is encrypted. Reports are stored in the patient's LGC folder, not on the inference workstation's file system. Purge temporary .soap.md files at the end of the day.
- Supervised updates
- Every model update (Whisper, LLM) triggers revalidation on a set of 10 anonymized test queries. Known regression bug = no deployment.
#Go further
This pipeline covers the everyday needs of a private practitioner or a small hospital department. Three natural extensions: industrialize multi-workstation deployment with Docker and a CPS-authentication reverse proxy, add a local RAG over HAS recommendations to assist with Assessment, and harden network security all the way to a complete air gap.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.