Local LLMs and GDPR: private-data compliance in enterprise
Sending customer data, contracts, or HR files to ChatGPT or Claude raises a simple question: where does the data go, who processes it, and under which jurisdiction? For a DPO, the answer is rarely compatible with the GDPR without contractual contortions. A local LLM — Ollama, vLLM, LM Studio — makes the question disappear: the data never leaves the company's infrastructure. This guide explains why a local GDPR-compliant enterprise LLM has become the default stack for DPOs in 2026, what changes with the AI Act, which fully takes effect in August, and how to audit your deployment.
#Why a local LLM is GDPR-native
The GDPR does not say "you are not allowed to use AI." It says that all processing of personal data must have a legal basis, be documented, minimized, and secured, and that transfers outside the EU must be governed. The problem with a cloud LLM is not AI—it is processing by a third-party processor, often American, with a complex contractual chain and transfer risks under Chapter V of the GDPR.
A locally run LLM reverses the equation. The model weights are downloaded once from Hugging Face or Ollama, then inference runs 100% on your hardware. No prompt is sent over the internet. No response is logged by a third party. The “processing” remains internal, under your effective control.
- No transfer outside the EU
- Article 44 and Chapter V of the GDPR: no data leaves your servers, so no Standard Contractual Clauses, no Transfer Impact Assessment, and no questions about the U.S. Cloud Act.
- No processor within the meaning of Article 28
- No outsourcing agreement to negotiate, no schedule 7 to update with every change to the vendor's product, and no vendor that unilaterally changes its standard terms and conditions.
- Native minimization
- Article 5.1.c: you cannot accidentally send too much data to a third party, because there is no third party. Minimization becomes an architectural property, not a policy that must be enforced.
- True deletion
- Article 17 (right to erasure): deleting a trace from your own database is trivial. Asking OpenAI to delete a prompt sent 6 months ago is a contractual procedure, not a technical guarantee.
#The legal framework in May 2026: GDPR + AI Act
Deploy local AI at work: privacy, compliance, multi-user architecture, costs, the one-page memo for leadership.
- Lifetime online access
- PDF + files
- Lifetime updates
Two texts now overlap in Europe to regulate the use of LLMs in businesses. The GDPR covers personal data, while the AI Act covers AI systems as such—regardless of whether they process personal data.
#GDPR: what applies to LLMs
- Legal basis (Art. 6)
- Legitimate interest, performance of a contract, consent: every LLM processing activity must be tied to one of these. A local LLM does not eliminate this obligation, but it simplifies the documentation.
- Information for individuals (Arts. 13-14)
- Your privacy policies must mention the use of an LLM, even a local one. You don't need to name the model, but you do need to state the purpose.
- DPIA (Art. 35)
- A Data Protection Impact Assessment is required for high-risk processing. With a cloud LLM, the DPIA must cover the processor; locally, it is limited to your own infrastructure.
- Security (Art. 32)
- Encryption of disks containing models and logs, access control for inference servers, logging. Standard for internal infrastructure.
#AI Act: the August 2026 deadline
The AI Act (EU Regulation 2024/1689) entered into force on 1ᵉʳ August 2024. Its provisions apply in stages. The most significant stage for general-purpose LLMs — the obligations for general-purpose AI models (GPAI) — took effect on 2 August 2025. The obligations for high-risk AI systems apply on 2 August 2026, approximately three months from now when you read this guide.
- GPAI (foundation models)
- The obligations apply to providers (OpenAI, Anthropic, Mistral, Meta...). If you use Mistral, Qwen or Granite locally, the publisher is already responsible for providing compliant technical documentation.
- High-risk systems
- Annex III: HR, credit scoring, education, essential public services, critical infrastructure. If your LLM is integrated into one of these workflows, you are a “deployer” under the AI Act and have specific obligations.
- Transparency
- All AI-generated content must be identifiable as such. Any system that interacts with a person must tell them so—including an internal chatbot.
- Penalties
- Up to €35M or 7% of global revenue for the most serious violations (prohibited use). For GPAI obligations, up to €15M or 3% of revenue.
#The concrete risks with ChatGPT, Claude, or Gemini
These providers now offer enterprise plans (ChatGPT Enterprise, Claude for Work, Gemini for Workspace) that contractually prohibit using data for training and sometimes provide European hosting. This is better than a consumer API. It doesn't solve everything.
- The US CLOUD Act
- An entity governed by US law (OpenAI Inc., Anthropic PBC, Google LLC) remains legally required to cooperate with US authorities, even for data stored in the EU. The CJEU reaffirmed this in Schrems II in 2020.
- The DPF on borrowed time
- The EU–US Data Privacy Framework (July 2023), which underpins most transfers, is being challenged before the CJEU. A Schrems III ruling could invalidate thousands of DPIAs overnight.
- Model opacity
- You don’t know exactly what the model saw during training, or how moderation filters log your prompts. On the ChatGPT side, “abuse” logs are retained for at least 30 days even in Enterprise.
- Shadow IT
- The practical #1 risk is not the contract; it's an employee copying and pasting an HR file into the free version of ChatGPT. No written policy can withstand a browser tab.
#The recommended stack for DPO in 2026
There isn't a single stack, but rather a family of proven combinations that French IT departments and public-sector IT departments have been deploying for 18 months. Three typical profiles cover 90% of needs.
#Profile 1 — Small team, individual workstations
- Hardware
- Workstations equipped with a RTX 4070 12 GB, RTX 4080 16 GB, or Mac M4 Pro 24-48 GB. No central server.
- Software
- Ollama (local daemon on each workstation) + LM Studio or Open WebUI as the interface. No data leaves the workstation.
- Recommended model
- Mistral Small 24B Q4 (14 GB, native French, excellent general-purpose model) or Qwen 3.8 27B at full quality if VRAM ≥ 40 GB (262k ctx, vision, Apache 2.0).
- Target
- Law firms, accounting firms, small-business HR teams, and journalists. Any organization with fewer than 30 people and individual use.
#Profile 2 — Internal inference server
- Hardware
- Dedicated GPU server: RTX 4090 24 GB, A6000 48 GB, or 2× RTX 3090 24 GB. Hosted in your IT environment or with a sovereign cloud provider (OVH, Scaleway, Outscale).
- Software
- vLLM or Ollama exposing an OpenAI-compatible API on the internal network. Open WebUI or LibreChat as the frontend, behind the company's IdP (Keycloak, Azure AD).
- Recommended model
- Mistral Small 24B, Qwen 3.6 35B-A3B or Granite 4.2 30B depending on VRAM (MoE models such as Qwen 3.6 35B-A3B activate only 3B parameters, so they are fast even on a 24 GB GPU). BGE-M3 or Solon embedding model for document RAG.
- Target
- Mid-sized companies, in-house legal departments, data teams of 30–500 people. Enables resource sharing and centralized access control.
#Profile 3 — Air gap for sensitive data
- Hardware
- Server physically isolated from the public network. LUKS-encrypted disks. Models downloaded via physical media on a transfer workstation.
- Software
- vLLM compiled locally, llama.cpp from source. No public container, no Docker image pulled on the fly.
- Recommended model
- Validated permissively licensed models (Mistral, Qwen, Granite, or Gemma 4 under Apache 2.0, with an audit of the weights). Ideally, a model for which you have archived a local copy of the tarball.
- Target
- Healthcare (DMP, medical reports), defense, highly sensitive trade secrets, OIV/OSE. Anything that would fall under an AIPD with high residual risk in the cloud.
#Implementation: the 5 foundational steps
- 01Map use casesBefore choosing a model, list the actual use cases: writing, translation, contract summarization, log analysis, and Tier 1 support. For each case, note the sensitivity of the data processed (public, internal, confidential, or professional secret). This mapping becomes the usage annex of your AIPD.
- 02Choose the profile and modelBased on the mapping and hardware inventory, select one of the three profiles above. For the model, start with Mistral Small 24B in Q4_K_M if you have 16-24 GB of VRAM—it’s the best-balanced tradeoff for French in 2026. Q4_K_M remains the default recommended quantization (quality loss < 2% vs FP16).
- 03Install the inference serverOllama listens on http://localhost:11434 by default. To expose it on the internal network, set OLLAMA_HOST=0.0.0.0:11434 and place the API behind a reverse proxy (Caddy, Traefik) that handles authentication. Enable access logs and retain them for 6 months for traceability.
- 04Document it in the registryCreate or update the processing record “Internal generative AI assistance” in your record of processing activities (Article 30 GDPR). Purposes, data categories, retention periods, and technical measures. The fact that it is local must be explicitly included.
- 05Train and communicateAn AI usage policy signed by every employee that states: (1) exclusive use of the internal instance, (2) prohibition of unapproved cloud tools, (3) data types allowed for each use case. Attach it to the employee handbook after consulting the CSE.
#Compliance audit checklist
This checklist covers the points that a GDPR–AI Act audit would verify for a local LLM deployment in an enterprise. Print it, check the items off, and archive it.
#Governance and documentation
- Art. 30 register
- Up-to-date “Internal Generative AI” processing record, including purposes, data, retention periods, recipients (internal only), and technical measures.
- AIPD
- Completed if the processing is high-risk (HR, healthcare, profiling). Updated whenever there is a major change to the model or use case.
- AI usage policy
- Distributed, signed, and incorporated into the internal regulations after consultation with the CSE.
- Use-case mapping
- Up-to-date list of validated use cases and authorized data for each case.
- Appointment of the AI lead
- An identified person (DPO, CISO, or IT department, depending on the organization), with a formal mandate.
#Technical security
- Encryption at rest
- Encrypted inference server disk(s) (LUKS, BitLocker, FileVault). Includes downloaded models and any prompt caches.
- Authentication
- API access filtered by OIDC or mTLS. No Ollama instance is accessible without authentication, even internally.
- Logging
- Access logs (who, when, model queried) retained for 6 to 12 months. No prompt-content logging except in explicitly documented use cases.
- Network isolation
- The inference server has no outbound internet access in production. For sensitive deployments, use a complete air gap.
- Backup and recovery
- Documented disaster recovery: models can be reinstalled from an internal archive, without immediate dependence on Hugging Face.
#AI Act compliance
- System classification
- You have determined that it is limited-, high-, or minimal-risk use. If it is high risk (Annex III), a specific compliance file exists.
- User information
- Each interface displays “AI-generated content” or an equivalent disclaimer. Article 50 of the AI Act, applicable August 2026.
- Model traceability
- Exact version of the model used, source of the weights, download date, archived SHA256 hashes. This lets you answer an audit question such as “which model generated this response on this date?”
- Human oversight
- For high-risk use cases, a documented human-review procedure before any decision affecting a person (hiring, scoring, disciplinary action).
#Common pitfalls
- “Local” but with telemetry
- Some interfaces (LM Studio in older versions, certain VSCode plugins) send telemetry. Verify with a sniffer (Wireshark, Little Snitch) that nothing is sent. Our privacy checklist guide covers this step.
- Models with ambiguous licenses
- Codestral 22B, for example, is under a non-production license: its use is prohibited in businesses. Llama, on the other hand, retains a community license that restricts use by very large companies (> 700 M MAU). Check the license before deploying to production—prefer Apache 2.0 models (Mistral Small, Qwen 3.5/3.8, Granite 4.2, Gemma 4)—especially if you are a large organization or a publisher.
- Model / fine-tuning confusion
- Fine-tuning on internal HR data creates a new processing activity, with its own DPIA. The fine-tuned model may reveal training data (memorization). For sensitive data, prefer RAG over fine-tuning.
- Underestimating traceability
- If in 18 months someone asks you, "show me what the AI answered about my case in March," you must be able to respond. Think about audit logs from day 1, not after the first incident.
- Believing that a local LLM solves everything
- The local LLM addresses data transfer, not use. A local model used for automated HR scoring remains a high-risk system under the meaning of the AI Act. Local ≠ exempt.
#Go further
Once compliance is in place, there are three natural directions for expanding the deployment:
- Lock down technical confidentiality
- The privacy checklist details the network and system checks that complement the legal framework on the DPO side. Essential before production deployment.
- Build a RAG system over your internal documents
- To move from a generic assistant to a business tool, RAG (Retrieval-Augmented Generation) lets you query contracts, procedures, and internal databases without fine-tuning. The introductory guide to local RAG lays the groundwork.
- Choose the user interface
- Open WebUI covers 80% of enterprise needs: multi-user support, OIDC, built-in RAG, and logs. The dedicated guide describes Docker deployment behind a reverse proxy.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.