Best GDPR-compliant on-premises LLM for businesses in 2026
Choose one GDPR-compliant on-premises LLM has become a central concern for European technical leadership teams that want to industrialize document processing without exposing their data to a non-European SaaS provider. A properly deployed on-premises GDPR-compliant LLM ensures that prompts, outputs, and any logs never leave the perimeter controlled by the organization, greatly simplifying impact assessments (DPIAs) and the obligations under Article 32 of the regulation. This article reviews the open-weight models relevant in 2026, their hardware footprint, licensing, use cases, and operational considerations before production deployment.
Why choose an on-premises GDPR-compliant LLM instead of a managed API
Regulation (EU) 2016/679 (see the consolidated text on EUR-Lex) requires strict controls on transfers outside the EU and clear documentation of the legal bases for processing. An API hosted by a third-party provider systematically introduces an additional processor, standard contractual clauses (CCT) to maintain, and in some cases exposure to extraterritorial laws such as the CLOUD Act. Conversely, an on-premises deployment (or a private cloud under the company’s control) removes this chain of dependencies.
Three other motivations come up in field feedback:
- Contractual sovereignty : no change to the provider's terms of service can interrupt the service.
- Fine-tuning proficiency : internal business-specific corpora remain within the information system, avoiding the risk of leakage through embeddings.
- Predictable marginal cost : once the hardware has paid for itself, inference no longer depends on the price per million tokens.
The open-weights ecosystem, cataloged on hubs such as Hugging Face or via the vLLM inference servers, now covers nearly all business use cases without external dependencies. For a broader overview of architectures, also see our open-source LLM guide for business.
Which open-weight models to choose in 2026
Choosing a sovereign enterprise LLM depends on a tradeoff between raw quality, context size, license, and VRAM footprint. Here is a representative selection drawn from our full catalog, organized by use-case profile.
"dense data center" profile (>200 GB Q4 VRAM)
- DeepSeek V4 Pro 1.6T : 1,600 billion parameters, MIT license, 1 000 000-token context, estimated Q4 VRAM of ~960 GB. Target: massive document RAG over legal or regulatory archives.
- MiMo V2.5 Pro : 1020B, MIT, 1M context, ~595 GB Q4 VRAM. Good multilingual versatility.
- Mistral Large 3 675B : 675B, Apache 2.0, 256,000 context, ~405 GB Q4 VRAM. French advantage: Paris-based publisher, with hosting fully controllable within the EU.
- GLM-5.1 : 744B, MIT, 200,000-token context, ~445 GB Q4 VRAM.
- Llama 4 Maverick 400B : 400B, Llama 4 Community license (note the commercial-use clauses beyond 700M MAU), 1M context.
“Intermediate cluster” profile (60–200 GB VRAM Q4)
- Qwen 3 235B-A22B : 235B mixture-of-experts (22B active), Apache 2.0, 131,072 context. Excellent quality-to-inference-cost ratio thanks to MoE.
- MiniMax-M2.7 : 229B, Apache 2.0, 205,000-token context.
- Mixtral 8x22B Instruct : 141B (39B active), Apache 2.0, 64,000 context. A solid European reference for RAG workloads.
- Qwen 3.5 122B-A10B : 122B, Apache 2.0, 262,000-token context.
- gpt-oss 120B : 117B, Apache 2.0, 128,000 context.
- Llama 4 Scout 109B : 109B, Llama 4 Community license, 10M-token context (to be confirmed in practice on very long prompts).
“Single-GPU server” profile (20–60 GB VRAM Q4)
- Llama 3.3 70B Instruct : 70B, Llama 3.3 Community license, 128,000 context. Fits on two RTX 6000 Ada cards or one 80 GB H100 in Q4.
- Apertus 70B : 70B, Apache 2.0, 65,536 context. Model originally from Switzerland, interesting for organizations seeking a broadly defined European vendor.
- Mixtral 8x7B : 47B, Apache 2.0, 32 768 context. Still relevant for batch workloads that are not highly sensitive to the state of the art.
- Salamandra 40B Instruct : 40B, Apache 2.0, 8192 context. From the Barcelona Supercomputing Center, trained on a European multilingual corpus.
- Qwen 3 32B et Qwen 2.5 Coder 32B : 32B, Apache 2.0. Targets general-purpose RAG and internal coding assistants.
For a more detailed comparison in the 70B segment, see Llama 3.3 70B vs Mixtral 8x22B.
VRAM footprint, quantization, and throughput
The VRAM estimate depends on the chosen quantization. As a first approximation, for a dense model:
- FP16 : ~2 bytes per parameter
- Q8 : ~1 byte per parameter
- Q5_K_M : ~0.7 bytes per parameter (estimated)
- Q4_K_M : ~0.55 to 0.60 bytes per parameter
For MoE architectures such as Mixtral 8x22B Instruct or Qwen 3 235B-A22B, the required VRAM corresponds to storing all the experts, even though only a few are activated per token. Memory is the constraint, not compute. The Mixtral 8x7B paper on arXiv details this mechanism.
In terms of throughput, tokens/sec figures vary widely depending on the inference engine (vLLM, TensorRT-LLM, llama.cpp, SGLang), effective context length, and batching. Some observed orders of magnitude (to be confirmed on your workload):
- Llama 3.3 70B Q4 on an H100 80 GB with vLLM: ~35–50 tokens/sec in single-stream, several hundred in batch.
- Mixtral 8x7B Q4 on RTX 4090 24 GB with partial offload: ~20–30 tokens/sec (estimated).
- Qwen 3 30B-A3B Q4 on RTX 6000 Ada 48 GB: ~60–80 tokens/sec in single-stream (estimated), helped by the MoE architecture.
- DeepSeek R1 671B Q4 on an 8×H200 node: ~15–25 tokens/sec in single-stream (to be confirmed depending on the inference version).
To size a server precisely, the quelllm.fr configurator cross-references GPUs, RAM, and compatible models.
Licensing: what really changes in production
Une GDPR-compliant AI is not enough: you also need a license compatible with your business model. Three families dominate:
- Apache 2.0 (Mistral, Qwen, gpt-oss, Mixtral, Snowflake, MiniMax, IBM Granite, BSC Salamandra…): unrestricted commercial use, redistribution permitted, patent protection clause. This is the simplest license for enterprise deployment.
- MIT (DeepSeek V3.2, R1, V4 Pro, GLM-5.1, Ling 2.6, MiMo, Ring-1T, dots.llm1, Seed-OSS): even more permissive, but without an explicit patent clause.
- Llama Community (Meta): commercial use authorized except beyond 700 million monthly active users, with acceptable use clauses. Readable on the Llama site.
- Gemma (Google): commercial, but subject to a prohibited-use policy.
For organizations seeking maximum traceability, OLMo 3 32B Allen AI also publishes its training datasets, which facilitates certain audits (see the Allen AI blog). On the European side, Mistral Large 3 675B et Apertus 70B are the core options. More details on our page best French LLM.
Benchmarks and use cases for a self-hosted enterprise LLM
Public scores should be treated with caution: methodologies vary, and test-set contamination is now documented. A few reference points from Hugging Face model cards:
- MMLU (5-shot general knowledge): models >200B typically reach 85-89%. DeepSeek R1 671B, Qwen 3 235B-A22B et Mistral Large 3 675B fall within this range (to be confirmed based on the evaluation pipeline).
- HumanEval / MBPP (Python code): Qwen 2.5 Coder 32B et Qwen3-Coder-Next 80B-A3B are serious candidates for an internal coding assistant.
- AIME 2024/2025 (mathematical reasoning): DeepSeek R1 671B, QwQ 32B et Ring-1T are positioned on this "reasoning" segment.
- Long-context (RULER, LongBench): Llama 4 Scout 109B, Seed-OSS 36B Instruct (524,288 tokens) and MiMo V2.5 Pro stand out on extended context lengths.
Typical use cases for IT departments:
- Internal document RAG : 32–70B is more than enough. See best LLM for RAG and our on-premises RAG guide.
- Legal / compliance assistant : prioritize a context of ≥ 128,000 tokens and a strong multilingual model (Mistral Large 3, Qwen 3 235B).
- Code generation : Qwen 2.5/3 Coder, Laguna XS.2, DeepSeek R2 32B.
- Multimodal : Qwen 3 VL 235B-A22B, Molmo 72B, LLaVA-OneVision 72B for analyzing scanned documents.
- Light workloads and edge : Granite 4.0 H-Small 32B-A9B, Gemma 4 31B, Qwen 3 30B-A3B.
See also our comparison DeepSeek R1 vs. Llama 3.3 70B to weigh reasoning against hardware cost.
Target architecture: from POC to production
A robust on-premises deployment is organized around four layers:
- Hardware layer : GPU NVIDIA (H100/H200/B200, RTX 6000 Ada, L40S) or AMD MI300X alternatives. For workloads >400B, an 8×H100 80 GB or 8×H200 141 GB node is the realistic minimum.
- Inference layer : vLLM or SGLang for throughput, TensorRT-LLM for latency, llama.cpp for servers without a dedicated GPU. The vLLM documentation covers most of the models mentioned here.
- Orchestration layer : Kubernetes with the NVIDIA GPU Operator, KEDA for scaling, and a LiteLLM- or Envoy-type gateway to unify OpenAI-compatible APIs.
- Compliance layer : encrypted prompt logging with short retention, PII masking on input (Presidio, CNIL recommendations on AI), an up-to-date processing register, and documented DPIA.
For observability, tools such as Langfuse or OpenTelemetry let you trace RAG chains without sending logs to a SaaS platform. For a structured approach, see our LLM production deployment guide and the LLM security checklist.
FAQ
Q: Is an open-weights LLM downloaded from Hugging Face automatically GDPR-compliant?
No. The model’s license and where it runs are two separate questions. Downloading the weights of Mistral Large 3 or DeepSeek R1 and running them in a European data center under your control eliminates the transfer of user data outside the EU, but you remain responsible for processing it (Article 24 of the GDPR): DPIAs, retention periods, data-subject rights, and access security.
Q: Which quantization should you choose for an on-premises GDPR-compliant LLM without degrading quality?
For production, Q5_K_M or Q4_K_M offer the best quality/VRAM trade-off on models >30B according to community evaluations published on Hugging Face. Below that (Q3, Q2), the loss becomes measurable on reasoning tasks. For sensitive regulated use, Q8 or FP16 are still recommended if VRAM allows, especially for models under 70B.
Q: Can you fine-tune on personal data while remaining compliant?
Yes, provided you document the legal basis, the DPIA, and retention. Models under the Apache 2.0 license (Mistral, Qwen, gpt-oss, IBM Granite) or MIT (DeepSeek, GLM) explicitly allow commercial fine-tuning. LoRA and QLoRA let you keep the adapters separate from the base weights, making it easier to delete them on request from a data subject when the corpus warrants it.
Q: Mixtral 8x22B or Llama 3.3 70B to get started?
Mixtral 8x22B Instruct uses more VRAM (82 GB Q4 versus 40 GB) but benefits from an Apache 2.0 license that is easier to integrate legally than the Llama Community License. Llama 3.3 70B Instruct remains very strong in French and on general-purpose tasks. For a first internal project, Llama 3.3 70B is often more accessible in terms of hardware; for a redistributable product, Mixtral is simpler from a licensing standpoint.
Q: Do you need an H100 cluster for a self-hosted enterprise LLM?
Not necessarily. A 2×RTX 6000 Ada 48 GB server or an 80 GB H100 is enough to run Llama 3.3 70B or Qwen 3 32B in Q4 on team RAG workloads. A multi-GPU cluster becomes necessary for models >100B at full precision or >400B when quantized. The quelllm.fr configurator calculates the minimum viable hardware.
Q: Do Chinese-origin models (DeepSeek, Qwen, GLM, MiMo) pose a GDPR issue?
The GDPR concerns data processing, not the origin of the weights. A DeepSeek model running locally does not communicate with a third-party server—it is a parameter file. The requirements to check are the license (MIT for DeepSeek, Apache 2.0 for Qwen) and any internal industry-specific restrictions. Security review (static analysis of the weights, sandboxing) remains recommended, as it is for any external artifact.
Conclusion
Adopt a GDPR-compliant on-premises LLM in 2026 is no longer an exploratory project: the open-weights ecosystem covers every workload profile, from Qwen 3 30B-A3B on a single GPU to DeepSeek V4 Pro 1.6T on a dense cluster, with Apache 2.0 or MIT licenses suitable for production. The real challenge has shifted to hardware sizing and industrialization. To identify the model suited to your hardware and constraints, run the configurator or browse the full catalog of the 249 indexed models.