Intermediate 11 minCompliance

Local AI in the enterprise: GDPR, sovereignty, and deployment (2026)

Direct response

A local AI server in France is a GPU machine (in your facilities or at a European hosting provider) that runs an open-weight model such as Mistral Small and shares it with your teams through a web interface: prompts and documents remain under your control. For an SME, a pilot can run on a workstation with a 24 GB GPU. Local does not mean outside GDPR: access, logging, records of processing activities, and security remain your responsibility.

You want to deploy an LLM in an organization without sending your data to a cloud service. This guide answers the questions in the order they arise: where to place the server, which model to choose (and under which license), what hardware to use for how many users, and how to deploy, secure, and document it for the DPO. Prices are not included here: they change every week.

By Mohamed Meguedmi·Update 2026-09-30·Tested on Windows, macOS, and Linux

#Local AI server in France: what it’s for and who it’s for

A local AI server hosts a language model on infrastructure you control: your own premises, a rack at a hosting provider, or a GPU server rented in Europe. Employees access it through a web interface or internal API, using their own accounts. For a small or midsize business, the starting setup is modest: a GPU workstation with 24 GB, Ollama as the engine, Open WebUI as the interface, and a 24-billion-parameter model such as Mistral Small 3.2, which occupies about 14 GB in Q4 according to the QuelLLM catalog. You then decide, based on a real use case, whether to scale up.

Privacy
Contracts, client files, HR data, source code: prompts and files never leave for an AI provider.
Reduced compliance, not eliminated
There is no AI processor to reference and no transfer outside the EU to document for the model itself, but your obligations as the data controller remain.
Predictable cost
No subscription that grows with your headcount: the cost is the hardware, electricity, and administration time.
Independence
A downloaded open-weight model continues to work if a provider changes its prices or terms.

#GDPR and sovereignty: what running locally changes—and what it does not

The AI at Work Kit

Deploy local AI at work: privacy, compliance, multi-user architecture, costs, the one-page memo for leadership.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

The CNIL notes that the RGPD protects people’s data in three places: in training databases, in models that may have memorized it, and when models are used through prompts. Your local server addresses the third issue, prompts: they stay with you. It does not address the second: a language model may itself contain personal data, which concerns your choice of model and how you document it.

For transfers, the CNIL states that a transfer outside the European Union and the European Economic Area is possible only if a sufficient and appropriate level of protection is ensured. With a server on your premises, the issue does not arise. With a hosting provider, it depends on its structure and subsidiaries.

!
Local doesn't mean outside the scope of the GDPR
You remain responsible for the processing: legal basis, informing individuals, minimization, retention period for conversation histories, access management, and logging. Add the assistant to the processing register and, if the processing is high-risk, have the need for an impact assessment evaluated with your DPO.

#Privately hosted LLM in France: concrete options

Three architectures can host a private LLM, depending on the level of control you want. The table compares them on what really matters: where the data is, who administers it, and what you must demonstrate to an auditor.

Three ways to host a private LLM
OptionWhere the data isWho administers itKey consideration
On-premises serverAt home, on your internal networkVousThe easiest to defend before a DPO; initial investment and operation are your responsibility
GPU server rented from a European hostAt the hosting provider, in the selected regionThe host provides the hardware; you provide the softwareCheck the actual region, the subcontracting agreement, and the hosting provider’s structure
Powerful workstation or mini PCOn the workstationVousIdeal for a pilot; no redundancy, no high availability

#Hosted in France, in Europe, SecNumCloud: three different things

“Hosted in France” describes a location, not legal protection. Scaleway, for example, highlights guaranteed data residency in Europe and protection against extraterritorial laws for its GPU instances: these are commercial commitments to review in the contract, not facts this guide can verify for you. For the most sensitive data, ANSSI offers SecNumCloud qualification, which aims to protect sensitive data and processing against cybercriminal threats and the application of extraterritorial laws.

Two useful clarifications before you rely on it. First, ANSSI notes that this qualification says nothing about the security level of the services you deploy on the qualified offering: your application still needs to be secured. Second, it qualifies specific cloud offerings, not an entire host: consult the catalog of qualified products and services to verify that the GPU offering you’re targeting is listed.

i
Ask a hosting provider for these four answers in writing
Exact GPU server region; contracting company and affiliated companies with administrator access; subcontracting agreement compliant with Article 28 of the GDPR; whether the offering is covered by a SecNumCloud qualification.

#Which model for a French company, and under which license

The choice comes down to three criteria: French-language quality, required memory, and licensing. Licensing is the criterion people forget: a very capable model may be prohibited for commercial use. The QuelLLM catalog lists each model’s license; here are a few reference points for French or European models.

French or European models for business use (QuelLLM catalog)
ModelSizeQ4 memoryLicenseNote
Mistral Nemo 12B Instruct12B7 GBApache 2.0128,000-token context; a good starting point on a small card
Mistral Small 3.2 24B24B14 GBApache 2.0Multilingual, vision; balance of quality and cost for an SMB
EuroLLM 22B Instruct22,6B13 GBApache 2.0European model; 32,768-token context
Mistral Small 4119B (MoE)72 GBApache 2.0For a high-memory server only
Codestral 22B v0.122B13 GBMistral Non-Production LicenseNot for commercial use: verify before any deployment

The 32-billion-parameter model often cited as a “good compromise” occupies 19 to 20 GB in Q4, weights only: on a 24 GB card, it leaves almost no room for context or multiple simultaneous users. For a first deployment, a 24-billion-parameter model is a more realistic starting point. Test French on your own documents before deciding: the quality shown in a ranking does not predict quality on your domain-specific jargon.

#Budget and hardware: size for concurrent users

Proper sizing depends not on the number of employees, but on the number of people sending a request at the same time. According to the Ollama FAQ, a model processes one request at a time by default; the others wait in a queue, and required memory grows as the product of the number of parallel requests and the context length. Allowing four parallel requests with an 8,000-token context therefore allocates the equivalent of 32,000 tokens of context memory.

Sizing guidelines (to be validated with a test using your workloads)
SituationWhat to aim forSoftware path
Drives 1 to 5 people, occasional use24 GB GPU workstation or Mac with large unified memory; 12B to 24B modelOllama and Open WebUI
A team of a few dozen people, daily useDedicated GPU server with more memory; measure latency with real usersOllama with tuned parallelism, or a high-throughput server such as vLLM
Multiple sites, availability requiredTwo servers, one gateway, monitoringLiteLLM or an equivalent in front of multiple engines

These figures are architecture guidelines, not measurements: perceived latency depends on the model, quantization, context, and usage profile. Run a two- to four-week pilot with a specific use case, monitor queues and memory usage, then size your system.

#Internal chatbot and document RAG

The most cost-effective use is a chatbot connected to your document repository: RAG, where the assistant answers using your procedures, contracts, or technical documentation retrieved on the fly. Open WebUI, AnythingLLM, and PrivateGPT offer this out of the box. Watch the interface license: Open WebUI prohibits removing its branding beyond fifty users over a rolling thirty-day period, unless authorized or covered by an enterprise license, which directly affects an enterprise deployment with custom branding.

Second point: RAG does not replace access rights. If all employees query the same database, a document reserved for management must remain inaccessible to everyone else. Isolate collections by group and test with an account without permissions.

#Deployment security

The starting point is simple: according to Ollama's FAQ, it listens on 127.0.0.1, port 11434 by default. As long as you don't modify OLLAMA_HOST, it can only be reached from the machine itself. The danger appears when you expose it to the network so the interface on another server can reach it, with nothing in front of it.

Network isolation
Server on an internal network, access through the interface only; complete isolation (air gap) for the most sensitive data.
Access management
Authentication, roles, disabling open registration, request logging in the interface.
Encryption
Encrypt the disks containing the models, document databases, and conversation histories.
Updates
Track engine, interface, and model versions; document who is responsible.
Backup and purge
Decide how long to retain conversations and define the deletion procedure when a person requests it.

#Recommended solutions and deployment in seven steps

  1. 01
    Frame a use case
    Choose a specific use case (support, internal documentation) and the relevant data category before choosing the hardware.
  2. 02
    Choose hosting
    Local, a European host, or a workstation, depending on data sensitivity and the options table.
  3. 03
    Choose the model
    Check the license, test French on your documents, and keep the required memory in mind.
  4. 04
    Install the engine and interface
    Ollama and Open WebUI for a pilot, a high-throughput server for a heavier workload.
  5. 05
    Secure
    Accounts, roles, HTTPS behind a reverse proxy, the engine port not exposed, encrypted disks.
  6. 06
    Document
    Processing records, informing individuals, retention period, data-processing agreement if using a hosting provider.
  7. 07
    Measure with a driver
    Two to four weeks of real-world use, then a scale-up decision.

This stack is free software, but it is not compliant “by design”: compliance depends on how you use it, and on whether that use is documented and controlled. Keep a record of your choices; that's what an auditor will ask for.

Frequently asked questions
Does local AI comply with the GDPR?+
It simplifies compliance because your prompts and documents do not go to a subcontractor, but it does not guarantee compliance. You remain responsible for the processing: legal basis, information, minimization, retention period, access, and logging. Add the processing to the register and have your DPO assess the need for an impact assessment.
How do you install an LLM in an enterprise?+
Define a use case, choose the hosting environment, then select a model whose license permits commercial use. Install Ollama and an interface such as Open WebUI, enable accounts and logging, encrypt the disks, and document the processing. Test for two to four weeks before sizing for scale.
Can you host a private LLM in France?+
Yes: on your premises, with a European host that guarantees the region, or on a workstation for a pilot. Check the exact region, the subcontracting agreement, and, for sensitive data, whether the offering has ANSSI’s SecNumCloud qualification. “Hosted in France” is not sufficient on its own.
Which server for local AI in an enterprise?+
It depends on the number of simultaneous requests, not the number of employees. A pilot fits on a workstation with a 24 GB GPU and a 12- to 24-billion-parameter model. For a few dozen users, plan for more memory and measure latency during a pilot before buying.
Is there a private French LLM?+
Yes: the Mistral models (Nemo, Small) are French and licensed under Apache 2.0 in the QuelLLM catalog, and EuroLLM is a European model. Pay attention to each model’s license: Codestral 22B, for example, is under a non-commercial license. Test French on your own documents before choosing.
Should you expose Ollama on the company network?+
It's best not to expose it: Ollama listens on 127.0.0.1 by default. If the interface is on another server, open the port only to that server, with a firewall, and route users through the authenticated interface behind an HTTPS reverse proxy. Never expose the engine's port to the Internet.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.