Best open-source LLM for the legal sector 2026
Choose one Open-source legal LLM in 2026 addresses a specific business requirement: maintaining control over customer data, respecting professional secrecy, and deploying a permissively licensed model on internal infrastructure. This guide compares the most relevant open-weights models for French-speaking law firms, legal departments, and legal tech companies, with hardware specifications, verifiable licenses, and use cases suited to French law. Below, you will find a selection from the quelllm.fr catalog, an overview of the Mistral, DeepSeek, Qwen, and Llama families, VRAM recommendations by model size, followed by an FAQ covering licensing, GDPR, and privacy questions.
Why an open-source LLM for the legal sector
The law handles sensitive data: client files, agreements, procedural documents, unpublished case law. Hosting a model on your own servers (or on sovereign cloud) eliminates prompt transmission to a third party and makes compliance with the RGPD as well as the obligations of theEU AI Act gradually came into effect starting in 2024.
Three concrete benefits for a Law firm AI self-hosted:
- Privacy : no prompt leaves the firm's environment, securing professional confidentiality guaranteed by Article 66-5 of the 1971 law.
- Domain-specific fine-tuning : possible adaptation to an internal case-law corpus (appellate court decisions, legal scholarship, standard pleadings).
- Controlled costs : no per-token billing; costs are limited to GPU infrastructure and energy.
The models under Apache 2.0 license or MIT are preferable: they allow commercial use, modification, and redistribution without a mandatory share-alike clause. Community licenses (Llama, Gemma) remain usable but require contractual review — see the open-weights licensing guide.
Recommended models for legal writing in French
Legal French remains an area where few models excel zero-shot. European providers and large multilingual models deliver the best results.
Mistral Family — the French option
- Mistral Large 3 675B (Apache 2.0, Mistral AI) — 675B parameters, ~405 GB VRAM in Q4, 256,000-token context. Designed in France, it handles French legal syntax and accepts very long documents (briefs, pleadings, consolidated contracts). Profile: Mistral Large 3.
- Mistral Small 4 (Apache 2.0) — 119B, ~72 GB VRAM Q4, 256,000 context. A good compromise for a midsize firm with 2× A100 80GB or 4× RTX 6000 Ada. Spec sheet: Mistral Small 4.
- Mixtral 8x22B Instruct (Apache 2.0) — 141B (39B active in MoE), ~82 GB VRAM in Q4. Solution Mistral legal robust and already proven in several legaltech deployments. Fact sheet: Mixtral 8x22B.
- Mixtral 8x7B (Apache 2.0)—47B, ~26 GB Q4 VRAM. The most accessible option to get started (1× RTX 4090 + possible CPU offload). Specs: Mixtral 8x7B.
Mistral AI publishes its tokenizers and weights on HuggingFace, which simplifies internal auditing.
Qwen family — high-performing multilingual models
Alibaba releases a Qwen family under Apache 2.0, with multilingual performance that includes French.
- Qwen 3 235B-A22B (~142 GB VRAM Q4, ctx 131 072) — MoE architecture with 22B active parameters, suited to summarizing long legal rulings. Details: Qwen 3 235B.
- Qwen 3.5 122B-A10B (~73 GB VRAM Q4, ctx 262 000) — 10B active, reduced latency for interactive writing.
- Qwen 3 32B (~19 GB Q4 VRAM) — deployable on a single RTX 4090 or A6000. Details: Qwen 3 32B.
- Qwen 3 VL 235B-A22B — vision variant for OCR of scanned documents and table reading. Specs: Qwen 3 VL 235B.
View the comparison Mistral Large 3 vs Qwen 3 235B to balance native French-language capability and multilingual versatility.
DeepSeek family—long reasoning
DeepSeek stands out in structured reasoning, which is useful for resource analysis, legal qualification, and building arguments.
- DeepSeek V3.2 (MIT, 685B, ~410 GB VRAM Q4, ctx 128 000)—Profile: DeepSeek V3.2.
- DeepSeek R1 671B (MIT, ~400 GB Q4 VRAM) — reasoning model published with the R1 article on arXiv. Specs: DeepSeek R1 671B.
- DeepSeek R1 Distill 32B (~19 GB VRAM Q4) — distilled version for a single-GPU server. Specs: DeepSeek R1 Distill 32B.
- DeepSeek R2 32B (~19 GB Q4 VRAM, ctx 128 000)—2026 iteration with improved reasoning.
Hardware specifications by cabinet size
The required VRAM depends on the chosen quantization. Approximate rule of thumb (to be confirmed for the llama.cpp or vLLM implementation):
- Q4 (4 bits): ~0.60 GB per billion parameters
- Q5 : ~0.75 GB / B
- Q8 : ~1.20 GB / B
- FP16 : ~2.00 GB / B
Solo practice or small firm (1 to 5 lawyers)
€2,000–6,000 GPU budget, 1× RTX 4090 24 GB or 1× RTX 6000 Ada 48 GB:
- Qwen 3 30B-A3B (~19 GB Q4, 3B active in MoE — very low latency) — fiche
- Mixtral 8x7B Q4 (~26 GB, partial offload)
- Gemma 4 31B (Gemma license, ~18 GB Q4, ctx 256 000) — fiche
- Qwen 3.6 35B-A3B (~21 GB Q4)
See best LLM for 24 GB VRAM for the complete selection.
Mid-sized firm (10 to 50 lawyers)
2× A100 80GB or 4× L40S 48GB server:
- Mistral Small 4 Q4 ~72 GB
- Mixtral 8x22B Q4 ~82 GB
- Qwen 3.5 122B-A10B Q4 ~73 GB
- gpt-oss 120B (Apache 2.0, OpenAI) ~70 GB — fiche
Large legal department or legaltech
Cluster of 8× H100 80GB (640 GB aggregated VRAM) or 4× MI300X 192 GB:
- Mistral Large 3 675B Q4 ~405 GB
- DeepSeek V3.2 ~410 GB Q4
- Llama 4 Maverick 400B (~240 GB Q4, ctx 1,000,000) — fiche
For multi-GPU architectures and tensor parallelism, see the distributed inference guide and the vLLM documentation.
Relevant benchmarks for legal work
No public benchmark specifically evaluates French law, but several proxies can be used:
- MMLU (subcategories
professional_lawetinternational_law) — primarily covers U.S. law but serves as an indicator. Reference: MMLU paper. - LegalBench (HuggingFace project) — 162 legal tasks in English.
- MMLU-Pro — hardened version, useful for legal qualification in the reasoning chain.
Based on these evaluations (public figures to be confirmed depending on the versions):
- Mistral Large 3 : estimated MMLU ~86%
- DeepSeek V3.2 : MMLU ~88%, MMLU-Pro ~75%
- Qwen 3 235B-A22B : MMLU ~87%
- Llama 3.3 70B Instruct : MMLU ~82%, ctx 128,000 — fiche
To compare coding and reasoning (useful for automating legal ops and clause analysis), see best LLM for coding et best reasoning LLM.
Concrete use cases in a practice
- Procedural document summary : Mistral Large 3 or Qwen 3.5 122B-A10B, with a context of ≥ 200,000 tokens capable of ingesting an entire folder.
- First draft (conclusions, letters, consultations): Mistral Small 4 or Mixtral 8x22B, sufficient to produce a draft reviewed by the lawyer.
- RAG-assisted case law research : Qwen 3 32B pipeline + Légifrance/Doctrine vector database, deployable on a single server.
- Contract analysis and clause detection : DeepSeek R2 32B or QwQ 32B for logically breaking down commitments — QwQ 32B spec sheet.
- Automatic anonymization of decisions : Granite 4.0 H-Small 32B-A9B (Apache 2.0, IBM, ~19 GB Q4) — fiche.
- OCR and reading scanned documents : Qwen 3 VL 235B-A22B for document multimodality.
A sovereign European project deserves mention: Apertus 70B (Swiss AI, Apache 2.0, ~40 GB Q4) — fiche —trained with particular attention to European compliance and Switzerland’s official languages, including French.
Inference performance (tokens/sec)
Estimates for common configurations (to be confirmed depending on the runtime and batch size):
- Mixtral 8x7B Q4 on RTX 4090 via llama.cpp: ~40-60 tok/s in single-stream
- Qwen 3 32B Q4 on RTX 6000 Ada via vLLM: ~35–50 tok/s
- Mistral Small 4 Q4 on 2× A100 80GB via vLLM: ~25–40 tok/s
- Mistral Large 3 Q4 on 8× H100 80GB: ~15–30 tok/s, with much higher batched throughput
See the measurement methods for the llama.cpp collective for reproducible benchmark protocols.
FAQ
Q: Which license should you choose for commercial deployment in a practice?
Prefer Apache 2.0 (Mistral, Qwen, Mixtral, gpt-oss, Apertus) or MIT (DeepSeek). These licenses explicitly allow commercial use, modification, and redistribution without a mandatory share-alike clause. The community licenses Llama and Gemma can be used but impose user caps (Llama) or usage restrictions (Gemma) that should be reviewed with legal before deployment.
Q: Is an open-source LLM hosted internally GDPR-compliant?
Self-hosting eliminates data transfer to a third party, which already resolves a major part of the risks. Final compliance nevertheless depends on the impact assessment (AIPD), the processing register, and retention periods for prompts and completions. Self-hosted deployment makes compliance easier but does not cover everything — a DPO audit is still required.
Q: How much VRAM for Mistral Large 3?
With Q4 quantization, Mistral Large 3 675B requires approximately 405 GB of VRAM aggregate, equivalent to 6× H100 80GB or 3× MI300X 192GB. In Q8 (more precise), plan for ~810 GB. For nearly equivalent quality with a smaller footprint, Mixtral 8x22B (~82 GB Q4) or Mistral Small 4 (~72 GB Q4) are reasonable alternatives.
Q: Which model for French law with a single 24 GB GPU?
With 24 GB of VRAM (RTX 4090, RTX 3090), the best candidates are Qwen 3 30B-A3B (~19 GB Q4, low latency thanks to MoE), Mixtral 8x7B Q4 (~26 GB with light offloading), Gemma 4 31B (~18 GB Q4) or DeepSeek R2 32B (~19 GB Q4). See the detailed selection at best LLM for 24 GB VRAM.
Q: Can you fine-tune a model on your own case law?
Yes, via LoRA or QLoRA for Apache 2.0 or MIT models. Expect 24 to 80 GB of VRAM depending on the base model size and chosen LoRA rank. MoE models (Mixtral, Qwen MoE) require special care with routing. See the legal fine-tuning guide for best practices on annotated corpora.
Q: Should you prefer a dense or MoE model in a professional practice?
A model dense (Llama 3.3 70B, Qwen 3 32B) delivers consistent quality and simple RAG integration. A model MoE (Mixtral 8x22B, Qwen 3 235B-A22B) offers a better quality-to-inference-cost ratio because only the active experts consume compute, but it requires more total VRAM to load all experts. For interactive use (assisted writing), MoE is advantageous; for batch processing (large-scale analysis), dense models remain competitive.
Conclusion
Choosing a Open-source legal LLM in 2026 is shaped by three factors: permissive licensing (Apache 2.0 or MIT prioritized), French-language quality (Mistral leading, followed by Qwen and DeepSeek), and a realistic GPU budget (from the RTX 4090 to an H100 cluster). To refine your selection based on your available VRAM and exact use case, run the quelllm.fr configurator or browse the full catalog of the 249 indexed models.