Best open-source LLM for the legal sector 2026

Choose one Open-source legal LLM in 2026 addresses a specific business requirement: maintaining control over customer data, respecting professional secrecy, and deploying a permissively licensed model on internal infrastructure. This guide compares the most relevant open-weights models for French-speaking law firms, legal departments, and legal tech companies, with hardware specifications, verifiable licenses, and use cases suited to French law. Below, you will find a selection from the quelllm.fr catalog, an overview of the Mistral, DeepSeek, Qwen, and Llama families, VRAM recommendations by model size, followed by an FAQ covering licensing, GDPR, and privacy questions.

Why an open-source LLM for the legal sector

The law handles sensitive data: client files, agreements, procedural documents, unpublished case law. Hosting a model on your own servers (or on sovereign cloud) eliminates prompt transmission to a third party and makes compliance with the RGPD as well as the obligations of theEU AI Act gradually came into effect starting in 2024.

Three concrete benefits for a Law firm AI self-hosted:

The models under Apache 2.0 license or MIT are preferable: they allow commercial use, modification, and redistribution without a mandatory share-alike clause. Community licenses (Llama, Gemma) remain usable but require contractual review — see the open-weights licensing guide.

Recommended models for legal writing in French

Legal French remains an area where few models excel zero-shot. European providers and large multilingual models deliver the best results.

Mistral Family — the French option

Mistral AI publishes its tokenizers and weights on HuggingFace, which simplifies internal auditing.

Qwen family — high-performing multilingual models

Alibaba releases a Qwen family under Apache 2.0, with multilingual performance that includes French.

View the comparison Mistral Large 3 vs Qwen 3 235B to balance native French-language capability and multilingual versatility.

DeepSeek family—long reasoning

DeepSeek stands out in structured reasoning, which is useful for resource analysis, legal qualification, and building arguments.

Hardware specifications by cabinet size

The required VRAM depends on the chosen quantization. Approximate rule of thumb (to be confirmed for the llama.cpp or vLLM implementation):

Solo practice or small firm (1 to 5 lawyers)

€2,000–6,000 GPU budget, 1× RTX 4090 24 GB or 1× RTX 6000 Ada 48 GB:

See best LLM for 24 GB VRAM for the complete selection.

Mid-sized firm (10 to 50 lawyers)

2× A100 80GB or 4× L40S 48GB server:

Large legal department or legaltech

Cluster of 8× H100 80GB (640 GB aggregated VRAM) or 4× MI300X 192 GB:

For multi-GPU architectures and tensor parallelism, see the distributed inference guide and the vLLM documentation.

Relevant benchmarks for legal work

No public benchmark specifically evaluates French law, but several proxies can be used:

Based on these evaluations (public figures to be confirmed depending on the versions):

To compare coding and reasoning (useful for automating legal ops and clause analysis), see best LLM for coding et best reasoning LLM.

Concrete use cases in a practice

A sovereign European project deserves mention: Apertus 70B (Swiss AI, Apache 2.0, ~40 GB Q4) — fiche —trained with particular attention to European compliance and Switzerland’s official languages, including French.

Inference performance (tokens/sec)

Estimates for common configurations (to be confirmed depending on the runtime and batch size):

See the measurement methods for the llama.cpp collective for reproducible benchmark protocols.

FAQ

Q: Which license should you choose for commercial deployment in a practice?

Prefer Apache 2.0 (Mistral, Qwen, Mixtral, gpt-oss, Apertus) or MIT (DeepSeek). These licenses explicitly allow commercial use, modification, and redistribution without a mandatory share-alike clause. The community licenses Llama and Gemma can be used but impose user caps (Llama) or usage restrictions (Gemma) that should be reviewed with legal before deployment.

Q: Is an open-source LLM hosted internally GDPR-compliant?

Self-hosting eliminates data transfer to a third party, which already resolves a major part of the risks. Final compliance nevertheless depends on the impact assessment (AIPD), the processing register, and retention periods for prompts and completions. Self-hosted deployment makes compliance easier but does not cover everything — a DPO audit is still required.

Q: How much VRAM for Mistral Large 3?

With Q4 quantization, Mistral Large 3 675B requires approximately 405 GB of VRAM aggregate, equivalent to 6× H100 80GB or 3× MI300X 192GB. In Q8 (more precise), plan for ~810 GB. For nearly equivalent quality with a smaller footprint, Mixtral 8x22B (~82 GB Q4) or Mistral Small 4 (~72 GB Q4) are reasonable alternatives.

Q: Which model for French law with a single 24 GB GPU?

With 24 GB of VRAM (RTX 4090, RTX 3090), the best candidates are Qwen 3 30B-A3B (~19 GB Q4, low latency thanks to MoE), Mixtral 8x7B Q4 (~26 GB with light offloading), Gemma 4 31B (~18 GB Q4) or DeepSeek R2 32B (~19 GB Q4). See the detailed selection at best LLM for 24 GB VRAM.

Q: Can you fine-tune a model on your own case law?

Yes, via LoRA or QLoRA for Apache 2.0 or MIT models. Expect 24 to 80 GB of VRAM depending on the base model size and chosen LoRA rank. MoE models (Mixtral, Qwen MoE) require special care with routing. See the legal fine-tuning guide for best practices on annotated corpora.

Q: Should you prefer a dense or MoE model in a professional practice?

A model dense (Llama 3.3 70B, Qwen 3 32B) delivers consistent quality and simple RAG integration. A model MoE (Mixtral 8x22B, Qwen 3 235B-A22B) offers a better quality-to-inference-cost ratio because only the active experts consume compute, but it requires more total VRAM to load all experts. For interactive use (assisted writing), MoE is advantageous; for batch processing (large-scale analysis), dense models remain competitive.

Conclusion

Choosing a Open-source legal LLM in 2026 is shaped by three factors: permissive licensing (Apache 2.0 or MIT prioritized), French-language quality (Mistral leading, followed by Qwen and DeepSeek), and a realistic GPU budget (from the RTX 4090 to an H100 cluster). To refine your selection based on your available VRAM and exact use case, run the quelllm.fr configurator or browse the full catalog of the 249 indexed models.

Article published and updated on by Mohamed Meguedmi · Data source: /api/models.json · Content license: CC BY 4.0.

An error or update to report? Contribute.

Prices in euros (€) are French market prices including VAT, checked by QuelLLM. US prices differ: the Amazon buttons show the current US price.