Case-file summaries medical
To summarize a medical record with a local LLM, extract the facts with code (FHIR batch or CDA R2 document), number them, have it draft the summary using only those facts and citing each line's reference, verify programmatically that every reference and number exists in the source, and then have a physician review it. Local execution protects confidentiality, not against hallucinations: a 2025 study found 1.47% per sentence.
A medical record is accurate but fragmented, and language models produce fluent summaries that may contain serious errors. This guide describes a local pipeline in which code extracts the facts, the model writes them up with citations, and mechanical verification followed by a doctor validates the result. It also highlights the framework to check: medical confidentiality, health-data hosting, and the tool's purpose.
#The problem: exact but unreadable folders
A practitioner taking over a patient’s care must review reports, lab results, prescriptions, and correspondence, often produced by multiple software systems. A local model can prepare a one-page handoff summary. The method that works reliably in a medical context is always the same: extract structured data with deterministic code, have the summary written from those facts alone, cite the identifier of the original fact on every line, check the summary programmatically, and then have a physician validate it.
One terminology point avoids a common misunderstanding. In France, exchanged and shared healthcare documents, particularly through the Dossier Médical Partagé, follow the interoperability framework for health information systems (CI-SIS) of the Agence du numérique en santé, whose sections describe documents in CDA R2 format, an HL7 standard based on XML. HL7 version 2 is another standard, based on messages with fields delimited by vertical bars, used for transport. Finally, FHIR, in JSON, is the most recent and easiest-to-process exchange standard, encountered primarily in exchanges between applications. Your first task is therefore to find out what your software provides.
#Framework: medical confidentiality, hosting, and responsibility
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
Three questions arise before writing any code. The first is medical confidentiality and health data protection, which require processing this data on infrastructure you control: local deployment is the practical implementation, but you must also encrypt the disks, restrict access, and keep logs. The second is health data hosting. Article L. 1111-8 of the French Public Health Code provides that anyone hosting personal health data on digital media must be certified for that purpose when doing so on behalf of a data controller or the patient.
This requirement does not apply in the way people often think. According to the French Digital Health Agency, as cited by a specialized insurer, hosting-provider certification is not a regulatory requirement when all of an organization's systems store only data belonging to its own patients, unless it provides hosting services on behalf of third parties. In other words, a practice running the tool in-house for its patients is not in the same situation as a vendor offering this service to other practices. Have your advisor or data protection officer confirm your situation.
The third question is the intended purpose. Software that produces information intended to inform a medical decision may fall under medical-device regulations, depending on its stated purpose. A case-summary handoff presented as reading assistance and validated by the physician does not have the same scope as a tool that suggests a treatment. This guide stays on the side of the former wording.
#The local stack
Three components are enough. Python to read the documents: the json module for a FHIR batch, lxml for a CDA document, and the hl7 library for possible version 2 messages, which describes itself as an HL7 v2.x message parser. Ollama to serve the model: Qwen 3.5 with 9 billion parameters weighs 6.6 GB and advertises 256,000 context tokens; Mistral Small 24B weighs 14 GB and advertises 32,000 tokens. Finally, the JSON schema from Ollama, to force the model to respond in a controllable format.
None of these models has been validated by its publishers for clinical use. Measure their medical French capabilities on your own records, using the protocol below. An 8 GB card is enough for the first; the second requires more video memory, especially with a long context.
#Read the documents: FHIR bundle and CDA document
For a FHIR bundle, which is a container for a collection of resources, the simplest approach is to read the JSON as-is and retrieve resources by type. This avoids depending on a library whose FHIR versions may vary. Each retained fact receives a short identifier (F1, F2...) and readable text; this is what will later make it possible to trace each sentence in the summary back to its source.
For a CDA document, the useful text is found in the structured body sections: each section has a title, a code, and a narrative text block. The code below extracts them with their titles, without including the document header, which contains the patient’s identity. The CDA namespaces are those of HL7 version 3.
#Which facts to remember, and with what caution
| Resource | What we read there | Common pitfall |
|---|---|---|
| Condition | Diagnosis, start date, status | A resolved or incorrect diagnosis can remain in the history |
| MedicationStatement | Medication, dosage, period | A stopped process may not be marked as such |
| Observation | Biology or measurement result, unit, date | Comparing values without their units or reference values |
| AllergyIntolerance | Substance, reaction, gravity | No reported intake does not mean no allergy |
| Procedure | Completed act and its date | Approximate or missing dates |
| DocumentReference | Attachment, often a report | Content encoded or returned as a link to retrieve separately |
The last column matters more than the others. A model reads the facts you give it and does not know what is missing: if a stopped treatment is not marked as such, it will appear as ongoing in the summary. The code must therefore include the status and date, and the prompt must ask it to flag facts with no date or uncertain status. Missing information is not information: the summary must say so.
#Write a summary that cites its sources, then verify it
The prompt provides numbered facts and requires every line of the summary to end with references to the facts used, such as [F3][F7]. It does not contain the patient’s name: age and sex are sufficient, and the identity remains in the practice-management software. The response format is free but structured into sections, making it readable on one page.
The next step is to enforce this in code. Two checks catch most serious errors: every cited reference must exist in the fact list, and every number in the summary must appear in the facts. A number appearing out of nowhere is the typical sign of a fabrication or transcription error, especially for a biological value or dosage.
#The system prompt
The “Points to verify” section is the most useful in practice. It turns ambiguities (treatment with no end date, a result without a unit, two contradictory values) into questions to ask the doctor instead of smoothing them into a fluent sentence.
#Evaluate reliability first
Trust cannot be decreed. A study published in 2025 in npj Digital Medicine measured language model errors when generating clinical notes: among 12,999 sentences annotated by clinicians, it found hallucinations in 1.47% of sentences and omissions in 3.45%, including 44% of hallucinations judged to be major—that is, potentially affecting diagnosis or care if not corrected. These figures concern a different task and different models: they are not yours. They show that even a low error rate per sentence remains concerning at the document scale.
A thirty-line summary with a 1.5% error rate per line, assuming independent errors, contains at least one error in about one-third of cases. This is an educational rule of thumb, not a forecast. Hence the following protocol, which replaces the “zero errors across 50 cases” criterion: that criterion proves very little, because zero observed errors across 50 cases is compatible with a true rate of about 6% at the 95% confidence level (the rule of three, 3 divided by 50).
- 01Build a sample of anonymized case filesUse varied records: patients with multiple conditions, numerous treatments, and older results. Remove identities and unnecessary data before any use outside the care setting.
- 02Have a doctor annotate itClassify each summary statement as correct, imprecise, omitted, or incorrect. Distinguish errors that could change a clinical decision.
- 03Count by category and severityThe number of major errors per case matters more than the average rate. Set a threshold with the responsible physician before starting.
- 04Measure omissions tooA summary that leaves out an allergy or treatment is more dangerous than an awkward sentence.
- 05Replay on every changeModel, prompt, or export format: every change requires rerunning the evaluation.
#Production deployment: traces, access, and scope
- Log every generation
- Keep the date, model version, supplied fact list, and generated summary in a protected space: that's what makes it possible to reconstruct what the doctor read.
- Control access
- The inference server must not be reachable from outside; disks are encrypted; accounts are named.
- Show sources
- The interface shows the summary with the original fact for each line. A doctor who can verify it with one click actually verifies it.
- Limit the scope
- Start with internal use within the firm or organization, with a small number of users, before expanding. A service offered to other organizations changes your hosting status.
- Plan for visible human validation
- The summary is signed or approved by the practitioner before entering the record; otherwise, it must not be included.
- Healthcare: consultation transcription
- Medical transcription and local LLMs: patient data
- Encrypt the model drive
- Privacy checklist
- Limit hallucinations in a local LLM
- Local LLMs and GDPR: private data in the enterprise
- Source: CI-SIS CDA-R2 transmission section (ANS)
- Source: healthcare data hosting (Relyens, according to ANS)
- Source: npj Digital Medicine study on clinical hallucinations
- Source: FHIR R4 Bundle resource
- Source: structured outputs from Ollama
Can you summarize a medical file with a local LLM?+
Does the DMP provide its documents in FHIR?+
Do you need an HDS-certified host for an internal tool at the firm?+
Which local model should you choose for medical text?+
Can an LLM make up a biology result?+
Is such a tool a medical device?+
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.