Local AI in the enterprise: GDPR, sovereignty, and deployment (2026)
A local AI server in France is a GPU machine (in your facilities or at a European hosting provider) that runs an open-weight model such as Mistral Small and shares it with your teams through a web interface: prompts and documents remain under your control. For an SME, a pilot can run on a workstation with a 24 GB GPU. Local does not mean outside GDPR: access, logging, records of processing activities, and security remain your responsibility.
You want to deploy an LLM in an organization without sending your data to a cloud service. This guide answers the questions in the order they arise: where to place the server, which model to choose (and under which license), what hardware to use for how many users, and how to deploy, secure, and document it for the DPO. Prices are not included here: they change every week.
#Local AI server in France: what it’s for and who it’s for
A local AI server hosts a language model on infrastructure you control: your own premises, a rack at a hosting provider, or a GPU server rented in Europe. Employees access it through a web interface or internal API, using their own accounts. For a small or midsize business, the starting setup is modest: a GPU workstation with 24 GB, Ollama as the engine, Open WebUI as the interface, and a 24-billion-parameter model such as Mistral Small 3.2, which occupies about 14 GB in Q4 according to the QuelLLM catalog. You then decide, based on a real use case, whether to scale up.
- Privacy
- Contracts, client files, HR data, source code: prompts and files never leave for an AI provider.
- Reduced compliance, not eliminated
- There is no AI processor to reference and no transfer outside the EU to document for the model itself, but your obligations as the data controller remain.
- Predictable cost
- No subscription that grows with your headcount: the cost is the hardware, electricity, and administration time.
- Independence
- A downloaded open-weight model continues to work if a provider changes its prices or terms.
#GDPR and sovereignty: what running locally changes—and what it does not
Deploy local AI at work: privacy, compliance, multi-user architecture, costs, the one-page memo for leadership.
- Lifetime online access
- PDF + files
- Lifetime updates
The CNIL notes that the RGPD protects people’s data in three places: in training databases, in models that may have memorized it, and when models are used through prompts. Your local server addresses the third issue, prompts: they stay with you. It does not address the second: a language model may itself contain personal data, which concerns your choice of model and how you document it.
For transfers, the CNIL states that a transfer outside the European Union and the European Economic Area is possible only if a sufficient and appropriate level of protection is ensured. With a server on your premises, the issue does not arise. With a hosting provider, it depends on its structure and subsidiaries.
#Privately hosted LLM in France: concrete options
Three architectures can host a private LLM, depending on the level of control you want. The table compares them on what really matters: where the data is, who administers it, and what you must demonstrate to an auditor.
| Option | Where the data is | Who administers it | Key consideration |
|---|---|---|---|
| On-premises server | At home, on your internal network | Vous | The easiest to defend before a DPO; initial investment and operation are your responsibility |
| GPU server rented from a European host | At the hosting provider, in the selected region | The host provides the hardware; you provide the software | Check the actual region, the subcontracting agreement, and the hosting provider’s structure |
| Powerful workstation or mini PC | On the workstation | Vous | Ideal for a pilot; no redundancy, no high availability |
#Hosted in France, in Europe, SecNumCloud: three different things
“Hosted in France” describes a location, not legal protection. Scaleway, for example, highlights guaranteed data residency in Europe and protection against extraterritorial laws for its GPU instances: these are commercial commitments to review in the contract, not facts this guide can verify for you. For the most sensitive data, ANSSI offers SecNumCloud qualification, which aims to protect sensitive data and processing against cybercriminal threats and the application of extraterritorial laws.
Two useful clarifications before you rely on it. First, ANSSI notes that this qualification says nothing about the security level of the services you deploy on the qualified offering: your application still needs to be secured. Second, it qualifies specific cloud offerings, not an entire host: consult the catalog of qualified products and services to verify that the GPU offering you’re targeting is listed.
#Which model for a French company, and under which license
The choice comes down to three criteria: French-language quality, required memory, and licensing. Licensing is the criterion people forget: a very capable model may be prohibited for commercial use. The QuelLLM catalog lists each model’s license; here are a few reference points for French or European models.
| Model | Size | Q4 memory | License | Note |
|---|---|---|---|---|
| Mistral Nemo 12B Instruct | 12B | 7 GB | Apache 2.0 | 128,000-token context; a good starting point on a small card |
| Mistral Small 3.2 24B | 24B | 14 GB | Apache 2.0 | Multilingual, vision; balance of quality and cost for an SMB |
| EuroLLM 22B Instruct | 22,6B | 13 GB | Apache 2.0 | European model; 32,768-token context |
| Mistral Small 4 | 119B (MoE) | 72 GB | Apache 2.0 | For a high-memory server only |
| Codestral 22B v0.1 | 22B | 13 GB | Mistral Non-Production License | Not for commercial use: verify before any deployment |
The 32-billion-parameter model often cited as a “good compromise” occupies 19 to 20 GB in Q4, weights only: on a 24 GB card, it leaves almost no room for context or multiple simultaneous users. For a first deployment, a 24-billion-parameter model is a more realistic starting point. Test French on your own documents before deciding: the quality shown in a ranking does not predict quality on your domain-specific jargon.
#Budget and hardware: size for concurrent users
Proper sizing depends not on the number of employees, but on the number of people sending a request at the same time. According to the Ollama FAQ, a model processes one request at a time by default; the others wait in a queue, and required memory grows as the product of the number of parallel requests and the context length. Allowing four parallel requests with an 8,000-token context therefore allocates the equivalent of 32,000 tokens of context memory.
| Situation | What to aim for | Software path |
|---|---|---|
| Drives 1 to 5 people, occasional use | 24 GB GPU workstation or Mac with large unified memory; 12B to 24B model | Ollama and Open WebUI |
| A team of a few dozen people, daily use | Dedicated GPU server with more memory; measure latency with real users | Ollama with tuned parallelism, or a high-throughput server such as vLLM |
| Multiple sites, availability required | Two servers, one gateway, monitoring | LiteLLM or an equivalent in front of multiple engines |
These figures are architecture guidelines, not measurements: perceived latency depends on the model, quantization, context, and usage profile. Run a two- to four-week pilot with a specific use case, monitor queues and memory usage, then size your system.
#Internal chatbot and document RAG
The most cost-effective use is a chatbot connected to your document repository: RAG, where the assistant answers using your procedures, contracts, or technical documentation retrieved on the fly. Open WebUI, AnythingLLM, and PrivateGPT offer this out of the box. Watch the interface license: Open WebUI prohibits removing its branding beyond fifty users over a rolling thirty-day period, unless authorized or covered by an enterprise license, which directly affects an enterprise deployment with custom branding.
Second point: RAG does not replace access rights. If all employees query the same database, a document reserved for management must remain inaccessible to everyone else. Isolate collections by group and test with an account without permissions.
#Deployment security
The starting point is simple: according to Ollama's FAQ, it listens on 127.0.0.1, port 11434 by default. As long as you don't modify OLLAMA_HOST, it can only be reached from the machine itself. The danger appears when you expose it to the network so the interface on another server can reach it, with nothing in front of it.
- Network isolation
- Server on an internal network, access through the interface only; complete isolation (air gap) for the most sensitive data.
- Access management
- Authentication, roles, disabling open registration, request logging in the interface.
- Encryption
- Encrypt the disks containing the models, document databases, and conversation histories.
- Updates
- Track engine, interface, and model versions; document who is responsible.
- Backup and purge
- Decide how long to retain conversations and define the deletion procedure when a person requests it.
#Recommended solutions and deployment in seven steps
- 01Frame a use caseChoose a specific use case (support, internal documentation) and the relevant data category before choosing the hardware.
- 02Choose hostingLocal, a European host, or a workstation, depending on data sensitivity and the options table.
- 03Choose the modelCheck the license, test French on your documents, and keep the required memory in mind.
- 04Install the engine and interfaceOllama and Open WebUI for a pilot, a high-throughput server for a heavier workload.
- 05SecureAccounts, roles, HTTPS behind a reverse proxy, the engine port not exposed, encrypted disks.
- 06DocumentProcessing records, informing individuals, retention period, data-processing agreement if using a hosting provider.
- 07Measure with a driverTwo to four weeks of real-world use, then a scale-up decision.
This stack is free software, but it is not compliant “by design”: compliance depends on how you use it, and on whether that use is documented and controlled. Keep a record of your choices; that's what an auditor will ask for.
- Source: CNIL, transferring data outside the EU
- Source: ANSSI, SecNumCloud qualification
- Source: Scaleway, GPU instances
- Source: Ollama FAQ
Does local AI comply with the GDPR?+
How do you install an LLM in an enterprise?+
Can you host a private LLM in France?+
Which server for local AI in an enterprise?+
Is there a private French LLM?+
Should you expose Ollama on the company network?+
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.