Intermediate 12 minPublic sector

Local AI for municipalities and local authorities territoriales

A town hall processes residents' correspondence, draft deliberations, meeting minutes, and civil-status requests every day. This is personal data, sometimes sensitive, that the local authority is legally required to protect. AI for local authorities as sold by vendors almost always relies on an American cloud, creating a legal and political problem. This guide shows how to deploy an open-weight LLM on a local-authority server, what you can use it for from the first week, and what it really costs, from a town hall serving 2,000 residents to an intermunicipal authority with 200 employees.

By Marie L.·Update 2026-09-28·Tested on Windows, macOS, and Linux

#Why local AI in local government

A local government's problem is not access to AI. Its staff already have it: ChatGPT on their personal phones, Copilot in the office suite, an assistant in their messaging app. The problem is that these uses happen without a framework, with residents’ data copied into services over which the local government has no control of hosting, retention, or reuse. The DPO (data protection officer, mandatory for every public authority since 2018) has no visibility into this.

A self-hosted LLM reverses the logic. The model runs on a machine at the local authority's premises or with its usual hosting provider. No text leaves the network. The data controller remains the local authority, with no additional processor to contract, no transfer outside the European Union, and no clause to negotiate with a giant that does not negotiate. And the cost is a one-time hardware investment, not a per-user subscription that grows with usage.

This is not a universal solution. A local model with 8 to 30 billion parameters is less capable than the best proprietary models on open-ended tasks. But municipal tasks are rarely open-ended: rewording a letter, structuring a resolution, summarizing a 40-page report, answering a question about cemetery regulations. On these well-defined tasks, with a good prompt and the right documents, a model the size of Mistral Small or Qwen3 14B gets the job done.

i
What exactly are we talking about?
Self-hosted does not mean improvised. The stack described here (Ollama for serving the model, Open WebUI for the multi-user interface) is the same one used by small and midsize businesses and firms handling sensitive data. It runs on Linux, is backed up like any other server, and can be administered by a local-government IT technician.

#Use cases that work in city halls

The AI at Work Kit

Deploy local AI at work: privacy, compliance, multi-user architecture, costs, the one-page memo for leadership.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

The best way to make an AI project fail in local government is to promise a “citizen assistant” that answers everything on the city website. The best way to make it succeed is to start with repetitive internal tasks where staff lose time and errors are easy to detect. Here are the ones that work well with a local model.

Letters to residents
Reply to a request for a school exemption, acknowledgment of receipt of a complaint, decision notification. The agent provides the factual details and the meaning of the response, the model drafts it in an administrative register, and the agent reviews and signs it. The gain is measured on low-stakes letters, which make up the majority.
Deliberation projects
Turn a service memo or quote into a structured draft resolution (citations, recitals, operative provisions). The model follows the format well when given two or three existing resolutions as templates. The substance remains with the general secretariat and the legal counsel.
Reports and summaries
Summarize a report from a research firm, extract decisions from a committee meeting summary, or produce a synthesis of a public consultation with 300 contributions. A model with a 32 000-token context can absorb a document of around a hundred pages.
Home and standard
Help the front-desk agent answer quickly and accurately: waste collection center hours, documents required for a birth certificate, after-school program rates. This uses RAG (retrieval-augmented generation) based on the local authority's official documents, not the model's memory.
Public procurement
First review of a CCTP to identify inconsistencies, reformulation of analysis criteria, and help drafting the bid evaluation report. The model makes no decisions; it speeds up the reading.
Translation and accessibility
Translate a notice for non-French-speaking residents, or rewrite a letter in plain language (FALC, easy to read and understand). Recent models are strong in French and the major European languages.

Two cases to rule out from the start. The first is any automated individual decision (awarding social assistance, calculating a family-quota-based rate): the Code governing relations between the public and the administration requires users to be informed of algorithmic processing and its rules to be explained, while the European AI regulation classifies assessing access to essential public services as a high-risk system. The second is a public-facing conversational agent on the city’s website, which requires a level of maturity (testing, supervision, and drift management) that few local authorities have for their first project.

→
Start with a service, not the organization
Choose a willing department with an engine-service lead: urban planning, general administration, or human resources. Three to five agents, one use case, six weeks. You’ll get concrete feedback, an initial quantified assessment in hours saved, and internal ambassadors before expanding.

#Why the U.S. cloud is problematic for a public authority

The argument is not ideological; it is legal. A public authority is a data controller under the GDPR. When it uses an online AI service, the provider becomes a processor, and a contract compliant with Article 28 must govern the relationship: purposes, retention periods, security, subprocessors, and what happens to the data when the contract ends. With major consumer services, this contract is a standard-form agreement that the public authority cannot amend, and the terms often allow conversations to be retained and, depending on the plan, used to improve the service.

Transfers outside the EU
Hosting in the United States is a transfer to a third country. Today, it relies on the EU–U.S. Data Privacy Framework (adequacy decision of July 2023), which has already been challenged before European courts. The Schrems II precedent (invalidation of the Privacy Shield in 2020) shows that this foundation can disappear overnight, and the public authority must then justify its processing on another basis.
Cloud Act
The 2018 U.S. law allows U.S. authorities to require a provider subject to U.S. law to hand over data, wherever it is stored. Hosting in a European data center operated by a U.S. provider does not protect against this requirement. This is precisely what the ANSSI SecNumCloud framework aims to rule out.
"Cloud at the center" doctrine
The 2021 Prime Minister's circular, updated in 2023, requires government administrations to host sensitive data on SecNumCloud-qualified offerings protected against extraterritorial legislation. It does not directly apply to local authorities, but it sets the level of rigor that an elected official or DPO can refer to, and prefectures draw on it in their dealings with local authorities.
Impact analysis
Processing residents’ data with generative AI at the scale of a local government falls into the category of cases where a data protection impact assessment (Article 35 of the GDPR) is strongly recommended. It is much simpler to write when there is neither a transfer nor a processor.
Secrets and sensitive data
Social services (CCAS), municipal police, and civil registries handle health data, offenses, and family situations. These categories are excluded from the terms of use of most consumer AI services, and sending them to a non-contracted third party constitutes a data breach that must be reported to the CNIL.

With self-hosting, most of these issues disappear by design. Obligations remain, and they are real: keep the processing register up to date with this new tool, define conversation retention periods, inform agents, and govern usage with a charter. But these are obligations the organization already knows how to meet for its other software.

!
Self-hosting doesn't exempt you from GDPR
A server in the city hall’s server room is still processing personal data. The DPO must be involved from the planning stage, the tool must be entered in the register, and conversation logs must be subject to a defined retention period. The advantage of local hosting is that it makes these obligations straightforward, not that it eliminates them.

#Prerequisites and governance

Before buying hardware, three decisions to make in a meeting, not at a terminal.

Who is behind the project
A two-person team: a business owner (director general of services, secretary general) and a technical lead (IT manager, the municipality's IT provider, or the intermunicipal authority's shared service). Without a business owner, the project remains the IT department's toy.
Where the server runs
At the local authority, in the intermunicipal body's or department's data center, or with a French hosting provider. All three are compatible with this guide. For a small municipality, the shared IT service of the community of municipalities is often the best option: one server, ten town halls.
What data goes in
Decide in advance which documents will feed the RAG (regulations, published deliberations, internal procedures) and which will not (individual CCAS files, municipal police data). This list is the core of the processing-activities register entry.
Skills
A technician comfortable with Linux and Docker is enough for installation and operation. RAG requires a little more process (document preparation), not more code.
Network
The server must be reachable from the agents’ workstations (internal network or VPN) and must never be exposed directly to the internet. An authenticated reverse proxy is the minimum if access extends beyond the local network.

On the purchasing side, a machine costing under €40,000 before tax falls under contracts without advertising or prior competitive bidding, which covers all the configurations in this guide. You still need a request for competing quotes and a justification of the need in the file. Most of the configurations described below can in fact be purchased through the UGAP catalog or from the local authority's usual system builder.

#Size the server and budget

The question that determines everything is the size of the model you want to serve, and therefore the available video memory (VRAM). In Q4_K_M quantization, the most common type, a model with 7 to 8 billion parameters occupies about 5 GB, a 14B model about 9 GB, a 24 to 32B model between 15 and 19 GB, and a 70B model about 40 GB. You must add context memory—the documents currently being read—which grows quickly with long texts. The number of simultaneous agents matters less than you might think: in a city hall, ten staff members equipped with the system rarely generate more than two or three requests at once.

Three tiers cover nearly all local authorities. The price ranges below are rough estimates, to be confirmed by quote: they vary by supplier, memory, storage, and warranty.

Tier 1: municipality with fewer than 5,000 residents
A workstation or mini-server with a card offering 12 or 16 GB of VRAM (such as RTX 4070, RTX 4080 or equivalent), 32 GB of RAM, and 1 TB of SSD storage. Target models: Qwen3 8B or 14B, Gemma 3 12B. Indicative hardware budget: €1,500 to €2,500 before tax. Five to ten agents, drafting and summarization, RAG over the municipality's regulations.
Tier 2: medium-sized city or intermunicipal authority
A workstation with a 24 GB card (RTX 4090 or a professional equivalent), 64 GB of RAM, 2 TB of SSD storage, or a Mac Studio with 64 GB of unified memory. Target models: Mistral Small 24B, Qwen3 32B in Q4, gpt-oss 20B. Indicative budget: €3,500 to €6,000 before tax. Twenty to fifty agents, covering all the use cases in this guide.
Tier 3: large local government or shared departmental service
A rack-mount server with one or two professional cards with 48 GB (or more), redundant power, integrated into the existing server room. Target models: 70B in Q4, or several specialized models served in parallel. Indicative budget starting at €10,000 excluding tax, often €15,000 to €25,000 excluding tax with manufacturer support. Several hundred agents, several member local authorities.

These amounts are supplemented by electricity (a workstation with a consumer card uses between 300 and 500 W under load and very little at idle), a few days of installation and configuration by the technical lead or a contractor, and above all the time spent training staff, which is the project's real cost. For comparison, a subscription to a proprietary AI assistant for fifty agents generally exceeds the price of tier 2 in the first year alone, and renews every year.

→
The Mac is a serious option for small teams
A Mac mini or Mac Studio with 48 to 64 GB of unified memory can run a 32B model without a dedicated graphics card, with very low power consumption and simple administration. It is less performant than a RTX 4090 for throughput, but more than sufficient for about ten agents, and the hardware is easy to buy from public catalogs.

#Step-by-step deployment

The process below assumes a Linux server (Ubuntu or Debian) with a NVIDIA card and the drivers installed. The stack is Ollama for serving the models and Open WebUI for the interface with user accounts, history, and built-in RAG.

  1. 01
    Install Ollama and make it accessible on the network
    The official script installs Ollama as a systemd service. By default, it listens only on the machine itself (http://localhost:11434). To allow the interface, installed in a container, to connect to it, make it listen on all interfaces, while ensuring that the server firewall blocks port 11434 from outside.
  2. 02
    Download one or two models
    Start with a medium-sized model and a small, fast model. Use the first for writing and summarization, and the second for short tasks and welcoming users. The download is several gigabytes, so schedule it outside business hours if the city hall’s connection is limited.
  3. 03
    Install Open WebUI
    Open WebUI is installed in a Docker container and connects to Ollama. The first account created is an administrator. Then disable open registration and create the agents’ accounts manually or through the directory (LDAP or SSO are supported in the administration settings).
  4. 04
    Build the document knowledge base
    In the Open WebUI workspace, create one collection per topic (regulations, deliberations, HR procedures) and upload the PDFs and Word documents. Prefer clean, up-to-date documents without image scans. Then link each collection to a custom model with a system prompt that describes the expected role.
  5. 05
    Write system prompts by profession
    A custom “Administrative mail” model with the local authority’s editorial guidelines, a “Resolution” model with the expected structure and two examples, and a “Welcome” model connected to the collection of practical information. These models appear as ready-to-use tools for staff, who do not need to know how to write prompts.
  6. 06
    Secure and back up
    Reverse proxy with a TLS certificate in front of Open WebUI, access limited to the internal network or VPN, daily backup of Open WebUI's data volume (accounts, conversations, documents) and configuration. Add the tool to the records of processing activities with a conversation-retention period, and configure purging accordingly.
Terminal (Linux server)
# Installation d'Ollama (script officiel)
curl -fsSL https://ollama.com/install.sh | sh

# Écouter sur le réseau pour le conteneur Open WebUI
sudo systemctl edit ollama.service
# Ajouter dans le fichier ouvert :
# [Service]
# Environment="OLLAMA_HOST=0.0.0.0"
# Environment="OLLAMA_KEEP_ALIVE=24h"
sudo systemctl daemon-reload
sudo systemctl restart ollama

# Modèles de départ (Q4_K_M par défaut)
ollama pull qwen3:14b
ollama pull mistral-small3.2
ollama pull qwen3:8b

# Vérification
ollama list
curl http://localhost:11434/api/tags
Terminal (Open WebUI via Docker)
docker run -d -p 3000:8080 \
  --add-host=host.docker.internal:host-gateway \
  -v open-webui:/app/backend/data \
  --name open-webui --restart always \
  ghcr.io/open-webui/open-webui:main

# L'interface est ensuite accessible sur http://<ip-du-serveur>:3000
# Premier compte créé = administrateur
# Ensuite : Admin > Settings > désactiver « Enable New Sign Ups »

To avoid leaving port 11434 open to the entire network, restrict it to the server itself and the container using the firewall (ufw or nftables). The site's guide to securing a Ollama server covers this section as well as setting up the reverse proxy. If the organization already has a Proxmox hypervisor, a virtual machine with GPU passthrough is a good way to integrate this server into the existing infrastructure.

i
Model and interface licenses
Check each model’s license before adopting it in a public service: Mistral Small, Qwen3, and gpt-oss are under Apache 2.0; Gemma 3 and Llama are under licenses specific to their respective vendors and are more restrictive. Open WebUI also adopted a license in 2025 that requires its brand to remain in the interface beyond a certain number of users, unless there is a commercial agreement. Nothing blocking for a local government, but document it in the project file.

#Support agents without replacing them

Fear of replacement is the first reaction in a department when AI is announced. It’s legitimate, and slogans won’t address it. What works is explaining precisely what the tool does and doesn’t do, then demonstrating it on the agents’ actual work.

The model proposes; the agent decides
No document produced by the model is released without human review and signature. This rule appears in the usage charter, and each customized model's system prompt reiterates it at the end of the response. It protects the organization (accountability) and the agent (who remains the author).
Train on real-world cases
Two hours per service, using the week’s letters and case files, not generic examples. The goal is for each agent to leave with three situations where the tool saves them time, and two where they shouldn’t use it.
Explain hallucinations
A local model invents legal references, code provisions, and figures with the same confidence as a correct answer. Agents must know that every legal reference produced by the model must be verified on Légifrance, and that RAG reduces this risk for the local authority’s documents without eliminating it.
Make usage visible
An internal convention: include the note “written with the help of an assistant, reviewed by [agent]” in working documents, never in official acts. It makes the practice less intimidating and facilitates oversight.
Involve employee representatives
Introducing an AI tool affects how work is organized. Informing the territorial social committee (CST) before deployment prevents it from learning about it through rumors and turning the project into a conflict.
Measure without monitoring
Track the number of active users and qualitative feedback, not individual productivity. Conversation logs are for troubleshooting and compliance, not agent evaluation, and the charter should say so.

In organizations that have taken this approach, adoption is not a problem: staff readily adopt a tool that removes the tedious part of a letter. The area requiring attention is elsewhere: some staff rely too heavily on the model and stop proofreading. Cross-checking during the first few weeks and regularly reminding users about observed hallucinations are the best safeguards.

#Pitfalls and troubleshooting

Slow or stuttering responses
The model exceeds VRAM capacity, and part of it runs on the processor. Check the GPU/CPU split with ollama ps. Choose a smaller model, use more aggressive quantization, or reduce the context size requested by Open WebUI.
RAG gives irrelevant answers
Most often, the documents are poorly segmented (scanned PDFs, tables, complex layouts). Convert them to clean text before uploading, remove obsolete versions of regulations, and test each collection with ten questions whose answers you know.
The model does not know the local-territory vocabulary
It doesn't know what a CCAS, DETR, or DSP means in your context. Add the local authority's glossary to the system prompt or to a dedicated collection.
An agent pasted a social services file into a conversation
This isn't an escape from the community, but it is out-of-scope processing. Delete the conversation, remind users of the charter, and if necessary adjust model access rights by user group in Open WebUI.
An update that breaks everything
Pin the versions: image tag Open WebUI, Ollama version, exact model list. Test updates on a secondary machine or virtual machine before touching the agent server.
The server is reachable from outside
Regularly check that no port forwarding or firewall rule exposes 11434 or the Open WebUI port on the city hall’s public address. Thousands of open Ollama servers are continuously indexed on the internet.

#Go further

This guide establishes the framework and basic deployment. These site guides explain the components you will work with next:

Deploy an AI chatbot for your team on the intranet
Details on the Nginx reverse proxy, authentication, monitoring, and conversation backups for a Open WebUI multi-user deployment.
Secure your Ollama server
Check your exposure, add authentication, TLS, and sound network practices. Read this before opening access beyond the local network.
Local LLMs and GDPR: private-data compliance
The complete legal framework (GDPR, the European AI Act, CNIL recommendations) for documenting the impact assessment and entering it in the register.
Local RAG with Ollama without coding
To understand what happens when you add documents to a collection and improve the quality of the welcome page's answers.
Choose your quantization (Q4, Q5, Q8, FP16)
To choose between a larger Q4 model and a smaller Q8 model based on the VRAM of the server you purchase.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.

Prices in euros (€) are French market prices including VAT, as checked by BestLLMfor. US prices differ: the Amazon buttons show the current US price.