Best open-source LLM for the medical sector in 2026

Choose one Open-source medical LLM in 2026 means balancing patient data confidentiality, clinical accuracy, and GPU budget. The Open-source medical LLM Self-hosting remains the preferred approach for GDPR compliance and professional confidentiality, without exfiltrating records to a third-party API. This guide reviews the most relevant open-weights models for healthcare use cases (report summarization, ICD-10 coding, differential diagnosis assistance, literature search), with VRAM specifications, licenses, and verifiable benchmarks.

Why self-host a model for healthcare

Personal health data falls under Article 9 of the GDPR and requires HDS hosting in France. A cloud API, even if encrypted, exposes the practice or hospital to a transfer outside the EU and complicates auditing. Self-hosting an open-weights model on a local GPU or on-premises server eliminates this risk.

The second argument is traceability. A locally weighted model (frozen weights, versioned system prompt) produces reproducible outputs, unlike a proprietary endpoint whose weights may change without notice. For a LLM diagnostics when used as an aid in medical decision-making, this reproducibility is required to qualify the device under MDR Regulation 2017/745.

The project Google Health’s MedGemma, often cited as a benchmark, illustrates the approach: open weights, fine-tuning on public medical corpora (MIMIC, PubMed), transparent evaluation. For a broader overview of the available families, see our complete catalog of the 249 indexed models.

Model families suited to clinical use cases

No general-purpose model is officially approved as a medical device in 2026. The choice is based on the quality of medical reasoning measured by the MedQA, MedMCQA, PubMedQA, and MMLU-Medical benchmarks, weighed against the hardware constraint.

Reasoning models (differential diagnosis, case synthesis)

Solid general-purpose models for healthcare

Multimodal models for imaging and scanned reports

VRAM, quantization, and target hardware

Hardware sizing determines every deployment oflocal healthcare AI. The values below concern inference, excluding training batches.

For a solo workstation (private practice, physician-researcher) : - Qwen 3.6 35B-A3B — Q4 ~21 GB, fits on a RTX 5090 32 GB - Granite 4.0 H-Small 32B-A9B — Q4 ~19 GB, IBM Research’s Apache 2.0 license - Gemma 4 31B — Q4 ~18 GB, Gemma license (read before clinical deployment) - Seed-OSS 36B Instruct — ctx 524,288, useful for ingesting several folders

For a hospital department (1 to 2 DGX or H100 servers) : - Llama 4 Scout 109B — Q4 ~65 GB, 10 million-token context (to be confirmed in production) - Mistral Small 4 — Q4 ~72 GB - Qwen 3.5 122B-A10B — Q4 ~73 GB - Mixtral 8x22B Instruct — Q4 ~82 GB, a reliable choice for medical writing

For a university hospital or healthcare publisher (multi-node cluster) : - DeepSeek V3.2 — Q4 ~410 GB - GLM-5.1 — Q4 ~445 GB, ctx 200,000 - Mistral Large 3 675B — Q4 ~405 GB

Q5 estimates add ~25%, and FP16 doubles these figures. To refine your estimate, use the GPU configurator which accounts for the KV cache at full context.

Licensing: what works in clinical settings and what blocks deployment

The license determines whether you can integrate the model into a product billed to a healthcare facility.

For European organizations concerned about sovereignty, Mistral Large 3 675B, Apertus 70B (Swiss AI) and Salamandra 40B Instruct (Barcelona Supercomputing Center) offer weights that can be hosted entirely within the EU. Compare these options on the page Mistral vs Llama.

Medical benchmarks: what the numbers mean

The most widely used public medical benchmarks in 2026:

MedQA scores for general-purpose models at pass@1 (estimates from HuggingFace model cards and technical papers, to be confirmed for critical use):

For programming biomedical pipeline tools (FHIR parsing, DICOM image extraction scripts), Qwen 2.5 Coder 32B or Qwen3-Coder-Next 80B-A3B achieve HumanEval scores above 85%.

Warning : a high MedQA benchmark score does not amount to market authorization. A Class IIa or higher medical device requires clinical validation compliant with IEC 62304 and the MDR.

Concrete use cases and recommended pipeline

Operative report summary : Mistral Small 4 or Llama 3.3 70B, 128,000-token context. Structured prompt requesting a SOAP summary. See our healthcare fine-tuning guide.

CIM-10 and CCAM coding assistance : Granite 4.0 H-Small with an RAG database based on the official nomenclature. IBM publishes evaluations on its Granite hub.

PubMed literature search : Qwen 3 VL 235B-A22B to read figures, coupled with a vector index. See the medical RAG guide.

Differential diagnosis support : DeepSeek R1 671B or DeepSeek R2 32B for step-by-step reasoning. Always with human validation, as an assistant and not a replacement.

Teleconsultation and transcription : combining Whisper-large-v3 (OpenAI repo) upstream, then Mistral Large 3 for formatting.

To compare with other sectors, see best legal LLM or best LLM for coding.

FAQ

Q: Is MedGemma listed on quelllm.fr?

MedGemma is not yet listed in the catalog of 249 indexed models because specialized healthcare variants are tracked separately. For immediate deployment, Gemma 4 31B serves as the foundation under the Gemma license, to be fine-tuned on an internal corpus. The MedGemma family is documented by Google Health on HuggingFace and remains compatible with a standard vLLM server.

Q: What is the minimum VRAM for solo clinical use?

Allow 20 to 24 GB to run a 32B model in Q4 with a useful 32 000-token context. A RTX 5090 32 GB, a RTX 6000 Ada 48 GB, or a Mac Studio M3 Ultra 96 GB will work. Below 16 GB, you are limited to 7B-13B models, whose medical quality is considered insufficient for professional use.

Q: Can you deploy an open-source medical LLM in production without certification?

For internal use to assist with writing without influencing medical decisions, yes. As soon as a module influences a diagnosis, prescription, or patient triage, you fall under the scope of the MDR Regulation 2017/745. The ANSM documentation on digital medical devices specifies the thresholds.

Q: What is the difference between DeepSeek R1 and R2 for healthcare?

DeepSeek R1 671B remains the reference for long-form reasoning and encyclopedic depth, but requires ~400 GB of VRAM. DeepSeek R2 32B is a distilled version running on a single GPU, with around 90% of the performance on MedQA according to available technical reports (to be confirmed).

Q: Do Chinese models pose a risk to patient data?

Open weights (DeepSeek, Qwen, GLM, MiMo) run locally without outbound telemetry. The risk is not exfiltration but cultural bias concerning conditions, dosage, and protocols. An internal evaluation on French-speaking cases is required. Compare with Mistral vs DeepSeek.

Q: Which model should I use for medical transcription in French?

For raw transcription, Whisper-large-v3 remains the benchmark. For formatting the verbatim transcript as meeting minutes, Mistral Small 4 or Llama 3.3 70B produce natural French medical language. Add a system prompt specifying the relevant specialty.

Conclusion

Choosing a Open-source medical LLM in 2026 depends first on the available hardware and the applicable regulatory framework. For a law firm, DeepSeek R2 32B or Granite 4.0 H-Small are sufficient. For a university hospital, Mistral Large 3 675B or DeepSeek R1 671B cover complex cases. Refine your selection with the GPU configurator or browse the full catalog filtered by license and VRAM.

Article published and updated on by Mohamed Meguedmi · Data source: /api/models.json · Content license: CC BY 4.0.

An error or update to report? Contribute.