Intermediate 15 minCompliance

Local LLMs and GDPR: private-data compliance in enterprise

Sending customer data, contracts, or HR files to ChatGPT or Claude raises a simple question: where does the data go, who processes it, and under which jurisdiction? For a DPO, the answer is rarely compatible with the GDPR without contractual contortions. A local LLM — Ollama, vLLM, LM Studio — makes the question disappear: the data never leaves the company's infrastructure. This guide explains why a local GDPR-compliant enterprise LLM has become the default stack for DPOs in 2026, what changes with the AI Act, which fully takes effect in August, and how to audit your deployment.

By Mohamed Meguedmi·Update 2026-08-27·Tested on Windows, macOS, and Linux

#Why a local LLM is GDPR-native

The GDPR does not say "you are not allowed to use AI." It says that all processing of personal data must have a legal basis, be documented, minimized, and secured, and that transfers outside the EU must be governed. The problem with a cloud LLM is not AI—it is processing by a third-party processor, often American, with a complex contractual chain and transfer risks under Chapter V of the GDPR.

A locally run LLM reverses the equation. The model weights are downloaded once from Hugging Face or Ollama, then inference runs 100% on your hardware. No prompt is sent over the internet. No response is logged by a third party. The “processing” remains internal, under your effective control.

No transfer outside the EU
Article 44 and Chapter V of the GDPR: no data leaves your servers, so no Standard Contractual Clauses, no Transfer Impact Assessment, and no questions about the U.S. Cloud Act.
No processor within the meaning of Article 28
No outsourcing agreement to negotiate, no schedule 7 to update with every change to the vendor's product, and no vendor that unilaterally changes its standard terms and conditions.
Native minimization
Article 5.1.c: you cannot accidentally send too much data to a third party, because there is no third party. Minimization becomes an architectural property, not a policy that must be enforced.
True deletion
Article 17 (right to erasure): deleting a trace from your own database is trivial. Asking OpenAI to delete a prompt sent 6 months ago is a contractual procedure, not a technical guarantee.
i
CNIL recommendation
Since 2024, CNIL has regularly published guidance on AI and the GDPR. Its consistent position: favor on-premises or sovereign solutions for processing sensitive data, health data, HR data, or any professional secret.
The AI at Work Kit

Deploy local AI at work: privacy, compliance, multi-user architecture, costs, the one-page memo for leadership.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

Two texts now overlap in Europe to regulate the use of LLMs in businesses. The GDPR covers personal data, while the AI Act covers AI systems as such—regardless of whether they process personal data.

#GDPR: what applies to LLMs

Legal basis (Art. 6)
Legitimate interest, performance of a contract, consent: every LLM processing activity must be tied to one of these. A local LLM does not eliminate this obligation, but it simplifies the documentation.
Information for individuals (Arts. 13-14)
Your privacy policies must mention the use of an LLM, even a local one. You don't need to name the model, but you do need to state the purpose.
DPIA (Art. 35)
A Data Protection Impact Assessment is required for high-risk processing. With a cloud LLM, the DPIA must cover the processor; locally, it is limited to your own infrastructure.
Security (Art. 32)
Encryption of disks containing models and logs, access control for inference servers, logging. Standard for internal infrastructure.

#AI Act: the August 2026 deadline

The AI Act (EU Regulation 2024/1689) entered into force on 1ᵉʳ August 2024. Its provisions apply in stages. The most significant stage for general-purpose LLMs — the obligations for general-purpose AI models (GPAI) — took effect on 2 August 2025. The obligations for high-risk AI systems apply on 2 August 2026, approximately three months from now when you read this guide.

GPAI (foundation models)
The obligations apply to providers (OpenAI, Anthropic, Mistral, Meta...). If you use Mistral, Qwen or Granite locally, the publisher is already responsible for providing compliant technical documentation.
High-risk systems
Annex III: HR, credit scoring, education, essential public services, critical infrastructure. If your LLM is integrated into one of these workflows, you are a “deployer” under the AI Act and have specific obligations.
Transparency
All AI-generated content must be identifiable as such. Any system that interacts with a person must tell them so—including an internal chatbot.
Penalties
Up to €35M or 7% of global revenue for the most serious violations (prohibited use). For GPAI obligations, up to €15M or 3% of revenue.
!
Deployer vs. provider status
Running Qwen 3.6 35B-A3B locally does not make you a GPAI provider. You remain a “deployer.” However, if you fine-tune a model and make it available to other entities, you may move into the provider category with its own obligations.

#The concrete risks with ChatGPT, Claude, or Gemini

These providers now offer enterprise plans (ChatGPT Enterprise, Claude for Work, Gemini for Workspace) that contractually prohibit using data for training and sometimes provide European hosting. This is better than a consumer API. It doesn't solve everything.

The US CLOUD Act
An entity governed by US law (OpenAI Inc., Anthropic PBC, Google LLC) remains legally required to cooperate with US authorities, even for data stored in the EU. The CJEU reaffirmed this in Schrems II in 2020.
The DPF on borrowed time
The EU–US Data Privacy Framework (July 2023), which underpins most transfers, is being challenged before the CJEU. A Schrems III ruling could invalidate thousands of DPIAs overnight.
Model opacity
You don’t know exactly what the model saw during training, or how moderation filters log your prompts. On the ChatGPT side, “abuse” logs are retained for at least 30 days even in Enterprise.
Shadow IT
The practical #1 risk is not the contract; it's an employee copying and pasting an HR file into the free version of ChatGPT. No written policy can withstand a browser tab.
→
The argument that resonates with the executive committee
A local LLM eliminates both legal risk (transfers, AI Act sanctions) and shadow IT risk (employees finally have an internal alternative that works). It's rarely “cheaper than ChatGPT Enterprise,” but it's almost always “less risky and faster to audit.”

#The recommended stack for DPO in 2026

There isn't a single stack, but rather a family of proven combinations that French IT departments and public-sector IT departments have been deploying for 18 months. Three typical profiles cover 90% of needs.

#Profile 1 — Small team, individual workstations

Hardware
Workstations equipped with a RTX 4070 12 GB, RTX 4080 16 GB, or Mac M4 Pro 24-48 GB. No central server.
Software
Ollama (local daemon on each workstation) + LM Studio or Open WebUI as the interface. No data leaves the workstation.
Recommended model
Mistral Small 24B Q4 (14 GB, native French, excellent general-purpose model) or Qwen 3.8 27B at full quality if VRAM ≥ 40 GB (262k ctx, vision, Apache 2.0).
Target
Law firms, accounting firms, small-business HR teams, and journalists. Any organization with fewer than 30 people and individual use.

#Profile 2 — Internal inference server

Hardware
Dedicated GPU server: RTX 4090 24 GB, A6000 48 GB, or 2× RTX 3090 24 GB. Hosted in your IT environment or with a sovereign cloud provider (OVH, Scaleway, Outscale).
Software
vLLM or Ollama exposing an OpenAI-compatible API on the internal network. Open WebUI or LibreChat as the frontend, behind the company's IdP (Keycloak, Azure AD).
Recommended model
Mistral Small 24B, Qwen 3.6 35B-A3B or Granite 4.2 30B depending on VRAM (MoE models such as Qwen 3.6 35B-A3B activate only 3B parameters, so they are fast even on a 24 GB GPU). BGE-M3 or Solon embedding model for document RAG.
Target
Mid-sized companies, in-house legal departments, data teams of 30–500 people. Enables resource sharing and centralized access control.

#Profile 3 — Air gap for sensitive data

Hardware
Server physically isolated from the public network. LUKS-encrypted disks. Models downloaded via physical media on a transfer workstation.
Software
vLLM compiled locally, llama.cpp from source. No public container, no Docker image pulled on the fly.
Recommended model
Validated permissively licensed models (Mistral, Qwen, Granite, or Gemma 4 under Apache 2.0, with an audit of the weights). Ideally, a model for which you have archived a local copy of the tarball.
Target
Healthcare (DMP, medical reports), defense, highly sensitive trade secrets, OIV/OSE. Anything that would fall under an AIPD with high residual risk in the cloud.

#Implementation: the 5 foundational steps

  1. 01
    Map use cases
    Before choosing a model, list the actual use cases: writing, translation, contract summarization, log analysis, and Tier 1 support. For each case, note the sensitivity of the data processed (public, internal, confidential, or professional secret). This mapping becomes the usage annex of your AIPD.
  2. 02
    Choose the profile and model
    Based on the mapping and hardware inventory, select one of the three profiles above. For the model, start with Mistral Small 24B in Q4_K_M if you have 16-24 GB of VRAM—it’s the best-balanced tradeoff for French in 2026. Q4_K_M remains the default recommended quantization (quality loss < 2% vs FP16).
  3. 03
    Install the inference server
    Ollama listens on http://localhost:11434 by default. To expose it on the internal network, set OLLAMA_HOST=0.0.0.0:11434 and place the API behind a reverse proxy (Caddy, Traefik) that handles authentication. Enable access logs and retain them for 6 months for traceability.
  4. 04
    Document it in the registry
    Create or update the processing record “Internal generative AI assistance” in your record of processing activities (Article 30 GDPR). Purposes, data categories, retention periods, and technical measures. The fact that it is local must be explicitly included.
  5. 05
    Train and communicate
    An AI usage policy signed by every employee that states: (1) exclusive use of the internal instance, (2) prohibition of unapproved cloud tools, (3) data types allowed for each use case. Attach it to the employee handbook after consulting the CSE.
Ollama configuration exposed on the internal network
# Sur Linux (systemd)
sudo systemctl edit ollama.service

# Ajouter :
[Service]
Environment="OLLAMA_HOST=0.0.0.0:11434"
Environment="OLLAMA_ORIGINS=https://chat.interne.entreprise.fr"
Environment="OLLAMA_KEEP_ALIVE=24h"

# Recharger et redémarrer
sudo systemctl daemon-reload
sudo systemctl restart ollama

# Vérifier
curl http://serveur-ia.interne:11434/api/tags
!
Do not expose 0.0.0.0 without a reverse proxy
OLLAMA_HOST=0.0.0.0 exposes the API without any authentication. Anyone on the network can query your models, list conversations, and download them. Always put an authenticated reverse proxy (OIDC, mTLS, basic auth + IP filtering) in front of it. This is non-negotiable.

#Compliance audit checklist

This checklist covers the points that a GDPR–AI Act audit would verify for a local LLM deployment in an enterprise. Print it, check the items off, and archive it.

#Governance and documentation

Art. 30 register
Up-to-date “Internal Generative AI” processing record, including purposes, data, retention periods, recipients (internal only), and technical measures.
AIPD
Completed if the processing is high-risk (HR, healthcare, profiling). Updated whenever there is a major change to the model or use case.
AI usage policy
Distributed, signed, and incorporated into the internal regulations after consultation with the CSE.
Use-case mapping
Up-to-date list of validated use cases and authorized data for each case.
Appointment of the AI lead
An identified person (DPO, CISO, or IT department, depending on the organization), with a formal mandate.

#Technical security

Encryption at rest
Encrypted inference server disk(s) (LUKS, BitLocker, FileVault). Includes downloaded models and any prompt caches.
Authentication
API access filtered by OIDC or mTLS. No Ollama instance is accessible without authentication, even internally.
Logging
Access logs (who, when, model queried) retained for 6 to 12 months. No prompt-content logging except in explicitly documented use cases.
Network isolation
The inference server has no outbound internet access in production. For sensitive deployments, use a complete air gap.
Backup and recovery
Documented disaster recovery: models can be reinstalled from an internal archive, without immediate dependence on Hugging Face.

#AI Act compliance

System classification
You have determined that it is limited-, high-, or minimal-risk use. If it is high risk (Annex III), a specific compliance file exists.
User information
Each interface displays “AI-generated content” or an equivalent disclaimer. Article 50 of the AI Act, applicable August 2026.
Model traceability
Exact version of the model used, source of the weights, download date, archived SHA256 hashes. This lets you answer an audit question such as “which model generated this response on this date?”
Human oversight
For high-risk use cases, a documented human-review procedure before any decision affecting a person (hiring, scoring, disciplinary action).

#Common pitfalls

“Local” but with telemetry
Some interfaces (LM Studio in older versions, certain VSCode plugins) send telemetry. Verify with a sniffer (Wireshark, Little Snitch) that nothing is sent. Our privacy checklist guide covers this step.
Models with ambiguous licenses
Codestral 22B, for example, is under a non-production license: its use is prohibited in businesses. Llama, on the other hand, retains a community license that restricts use by very large companies (> 700 M MAU). Check the license before deploying to production—prefer Apache 2.0 models (Mistral Small, Qwen 3.5/3.8, Granite 4.2, Gemma 4)—especially if you are a large organization or a publisher.
Model / fine-tuning confusion
Fine-tuning on internal HR data creates a new processing activity, with its own DPIA. The fine-tuned model may reveal training data (memorization). For sensitive data, prefer RAG over fine-tuning.
Underestimating traceability
If in 18 months someone asks you, "show me what the AI answered about my case in March," you must be able to respond. Think about audit logs from day 1, not after the first incident.
Believing that a local LLM solves everything
The local LLM addresses data transfer, not use. A local model used for automated HR scoring remains a high-risk system under the meaning of the AI Act. Local ≠ exempt.
i
Summary of the CNIL position
Since 2024, CNIL has published several AI fact sheets (on the legal basis, data minimization, and AIPD-IA). It does not require local deployment but values it as an "appropriate technical measure" within the meaning of Article 32 of the GDPR, especially for sensitive data.

#Go further

Once compliance is in place, there are three natural directions for expanding the deployment:

Lock down technical confidentiality
The privacy checklist details the network and system checks that complement the legal framework on the DPO side. Essential before production deployment.
Build a RAG system over your internal documents
To move from a generic assistant to a business tool, RAG (Retrieval-Augmented Generation) lets you query contracts, procedures, and internal databases without fine-tuning. The introductory guide to local RAG lays the groundwork.
Choose the user interface
Open WebUI covers 80% of enterprise needs: multi-user support, OIDC, built-in RAG, and logs. The dedicated guide describes Docker deployment behind a reverse proxy.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.

Prices in euros (€) are French market prices including VAT, as checked by BestLLMfor. US prices differ: the Amazon buttons show the current US price.