Local AI for municipalities and local authorities territoriales
A town hall processes residents' correspondence, draft deliberations, meeting minutes, and civil-status requests every day. This is personal data, sometimes sensitive, that the local authority is legally required to protect. AI for local authorities as sold by vendors almost always relies on an American cloud, creating a legal and political problem. This guide shows how to deploy an open-weight LLM on a local-authority server, what you can use it for from the first week, and what it really costs, from a town hall serving 2,000 residents to an intermunicipal authority with 200 employees.
#Why local AI in local government
A local government's problem is not access to AI. Its staff already have it: ChatGPT on their personal phones, Copilot in the office suite, an assistant in their messaging app. The problem is that these uses happen without a framework, with residents’ data copied into services over which the local government has no control of hosting, retention, or reuse. The DPO (data protection officer, mandatory for every public authority since 2018) has no visibility into this.
A self-hosted LLM reverses the logic. The model runs on a machine at the local authority's premises or with its usual hosting provider. No text leaves the network. The data controller remains the local authority, with no additional processor to contract, no transfer outside the European Union, and no clause to negotiate with a giant that does not negotiate. And the cost is a one-time hardware investment, not a per-user subscription that grows with usage.
This is not a universal solution. A local model with 8 to 30 billion parameters is less capable than the best proprietary models on open-ended tasks. But municipal tasks are rarely open-ended: rewording a letter, structuring a resolution, summarizing a 40-page report, answering a question about cemetery regulations. On these well-defined tasks, with a good prompt and the right documents, a model the size of Mistral Small or Qwen3 14B gets the job done.
#Use cases that work in city halls
Deploy local AI at work: privacy, compliance, multi-user architecture, costs, the one-page memo for leadership.
- Lifetime online access
- PDF + files
- Lifetime updates
The best way to make an AI project fail in local government is to promise a “citizen assistant” that answers everything on the city website. The best way to make it succeed is to start with repetitive internal tasks where staff lose time and errors are easy to detect. Here are the ones that work well with a local model.
- Letters to residents
- Reply to a request for a school exemption, acknowledgment of receipt of a complaint, decision notification. The agent provides the factual details and the meaning of the response, the model drafts it in an administrative register, and the agent reviews and signs it. The gain is measured on low-stakes letters, which make up the majority.
- Deliberation projects
- Turn a service memo or quote into a structured draft resolution (citations, recitals, operative provisions). The model follows the format well when given two or three existing resolutions as templates. The substance remains with the general secretariat and the legal counsel.
- Reports and summaries
- Summarize a report from a research firm, extract decisions from a committee meeting summary, or produce a synthesis of a public consultation with 300 contributions. A model with a 32 000-token context can absorb a document of around a hundred pages.
- Home and standard
- Help the front-desk agent answer quickly and accurately: waste collection center hours, documents required for a birth certificate, after-school program rates. This uses RAG (retrieval-augmented generation) based on the local authority's official documents, not the model's memory.
- Public procurement
- First review of a CCTP to identify inconsistencies, reformulation of analysis criteria, and help drafting the bid evaluation report. The model makes no decisions; it speeds up the reading.
- Translation and accessibility
- Translate a notice for non-French-speaking residents, or rewrite a letter in plain language (FALC, easy to read and understand). Recent models are strong in French and the major European languages.
Two cases to rule out from the start. The first is any automated individual decision (awarding social assistance, calculating a family-quota-based rate): the Code governing relations between the public and the administration requires users to be informed of algorithmic processing and its rules to be explained, while the European AI regulation classifies assessing access to essential public services as a high-risk system. The second is a public-facing conversational agent on the city’s website, which requires a level of maturity (testing, supervision, and drift management) that few local authorities have for their first project.
#Why the U.S. cloud is problematic for a public authority
The argument is not ideological; it is legal. A public authority is a data controller under the GDPR. When it uses an online AI service, the provider becomes a processor, and a contract compliant with Article 28 must govern the relationship: purposes, retention periods, security, subprocessors, and what happens to the data when the contract ends. With major consumer services, this contract is a standard-form agreement that the public authority cannot amend, and the terms often allow conversations to be retained and, depending on the plan, used to improve the service.
- Transfers outside the EU
- Hosting in the United States is a transfer to a third country. Today, it relies on the EU–U.S. Data Privacy Framework (adequacy decision of July 2023), which has already been challenged before European courts. The Schrems II precedent (invalidation of the Privacy Shield in 2020) shows that this foundation can disappear overnight, and the public authority must then justify its processing on another basis.
- Cloud Act
- The 2018 U.S. law allows U.S. authorities to require a provider subject to U.S. law to hand over data, wherever it is stored. Hosting in a European data center operated by a U.S. provider does not protect against this requirement. This is precisely what the ANSSI SecNumCloud framework aims to rule out.
- "Cloud at the center" doctrine
- The 2021 Prime Minister's circular, updated in 2023, requires government administrations to host sensitive data on SecNumCloud-qualified offerings protected against extraterritorial legislation. It does not directly apply to local authorities, but it sets the level of rigor that an elected official or DPO can refer to, and prefectures draw on it in their dealings with local authorities.
- Impact analysis
- Processing residents’ data with generative AI at the scale of a local government falls into the category of cases where a data protection impact assessment (Article 35 of the GDPR) is strongly recommended. It is much simpler to write when there is neither a transfer nor a processor.
- Secrets and sensitive data
- Social services (CCAS), municipal police, and civil registries handle health data, offenses, and family situations. These categories are excluded from the terms of use of most consumer AI services, and sending them to a non-contracted third party constitutes a data breach that must be reported to the CNIL.
With self-hosting, most of these issues disappear by design. Obligations remain, and they are real: keep the processing register up to date with this new tool, define conversation retention periods, inform agents, and govern usage with a charter. But these are obligations the organization already knows how to meet for its other software.
#Prerequisites and governance
Before buying hardware, three decisions to make in a meeting, not at a terminal.
- Who is behind the project
- A two-person team: a business owner (director general of services, secretary general) and a technical lead (IT manager, the municipality's IT provider, or the intermunicipal authority's shared service). Without a business owner, the project remains the IT department's toy.
- Where the server runs
- At the local authority, in the intermunicipal body's or department's data center, or with a French hosting provider. All three are compatible with this guide. For a small municipality, the shared IT service of the community of municipalities is often the best option: one server, ten town halls.
- What data goes in
- Decide in advance which documents will feed the RAG (regulations, published deliberations, internal procedures) and which will not (individual CCAS files, municipal police data). This list is the core of the processing-activities register entry.
- Skills
- A technician comfortable with Linux and Docker is enough for installation and operation. RAG requires a little more process (document preparation), not more code.
- Network
- The server must be reachable from the agents’ workstations (internal network or VPN) and must never be exposed directly to the internet. An authenticated reverse proxy is the minimum if access extends beyond the local network.
On the purchasing side, a machine costing under €40,000 before tax falls under contracts without advertising or prior competitive bidding, which covers all the configurations in this guide. You still need a request for competing quotes and a justification of the need in the file. Most of the configurations described below can in fact be purchased through the UGAP catalog or from the local authority's usual system builder.
#Size the server and budget
The question that determines everything is the size of the model you want to serve, and therefore the available video memory (VRAM). In Q4_K_M quantization, the most common type, a model with 7 to 8 billion parameters occupies about 5 GB, a 14B model about 9 GB, a 24 to 32B model between 15 and 19 GB, and a 70B model about 40 GB. You must add context memory—the documents currently being read—which grows quickly with long texts. The number of simultaneous agents matters less than you might think: in a city hall, ten staff members equipped with the system rarely generate more than two or three requests at once.
Three tiers cover nearly all local authorities. The price ranges below are rough estimates, to be confirmed by quote: they vary by supplier, memory, storage, and warranty.
- Tier 1: municipality with fewer than 5,000 residents
- A workstation or mini-server with a card offering 12 or 16 GB of VRAM (such as RTX 4070, RTX 4080 or equivalent), 32 GB of RAM, and 1 TB of SSD storage. Target models: Qwen3 8B or 14B, Gemma 3 12B. Indicative hardware budget: €1,500 to €2,500 before tax. Five to ten agents, drafting and summarization, RAG over the municipality's regulations.
- Tier 2: medium-sized city or intermunicipal authority
- A workstation with a 24 GB card (RTX 4090 or a professional equivalent), 64 GB of RAM, 2 TB of SSD storage, or a Mac Studio with 64 GB of unified memory. Target models: Mistral Small 24B, Qwen3 32B in Q4, gpt-oss 20B. Indicative budget: €3,500 to €6,000 before tax. Twenty to fifty agents, covering all the use cases in this guide.
- Tier 3: large local government or shared departmental service
- A rack-mount server with one or two professional cards with 48 GB (or more), redundant power, integrated into the existing server room. Target models: 70B in Q4, or several specialized models served in parallel. Indicative budget starting at €10,000 excluding tax, often €15,000 to €25,000 excluding tax with manufacturer support. Several hundred agents, several member local authorities.
These amounts are supplemented by electricity (a workstation with a consumer card uses between 300 and 500 W under load and very little at idle), a few days of installation and configuration by the technical lead or a contractor, and above all the time spent training staff, which is the project's real cost. For comparison, a subscription to a proprietary AI assistant for fifty agents generally exceeds the price of tier 2 in the first year alone, and renews every year.
#Step-by-step deployment
The process below assumes a Linux server (Ubuntu or Debian) with a NVIDIA card and the drivers installed. The stack is Ollama for serving the models and Open WebUI for the interface with user accounts, history, and built-in RAG.
- 01Install Ollama and make it accessible on the networkThe official script installs Ollama as a systemd service. By default, it listens only on the machine itself (http://localhost:11434). To allow the interface, installed in a container, to connect to it, make it listen on all interfaces, while ensuring that the server firewall blocks port 11434 from outside.
- 02Download one or two modelsStart with a medium-sized model and a small, fast model. Use the first for writing and summarization, and the second for short tasks and welcoming users. The download is several gigabytes, so schedule it outside business hours if the city hall’s connection is limited.
- 03Install Open WebUIOpen WebUI is installed in a Docker container and connects to Ollama. The first account created is an administrator. Then disable open registration and create the agents’ accounts manually or through the directory (LDAP or SSO are supported in the administration settings).
- 04Build the document knowledge baseIn the Open WebUI workspace, create one collection per topic (regulations, deliberations, HR procedures) and upload the PDFs and Word documents. Prefer clean, up-to-date documents without image scans. Then link each collection to a custom model with a system prompt that describes the expected role.
- 05Write system prompts by professionA custom “Administrative mail” model with the local authority’s editorial guidelines, a “Resolution” model with the expected structure and two examples, and a “Welcome” model connected to the collection of practical information. These models appear as ready-to-use tools for staff, who do not need to know how to write prompts.
- 06Secure and back upReverse proxy with a TLS certificate in front of Open WebUI, access limited to the internal network or VPN, daily backup of Open WebUI's data volume (accounts, conversations, documents) and configuration. Add the tool to the records of processing activities with a conversation-retention period, and configure purging accordingly.
To avoid leaving port 11434 open to the entire network, restrict it to the server itself and the container using the firewall (ufw or nftables). The site's guide to securing a Ollama server covers this section as well as setting up the reverse proxy. If the organization already has a Proxmox hypervisor, a virtual machine with GPU passthrough is a good way to integrate this server into the existing infrastructure.
#Support agents without replacing them
Fear of replacement is the first reaction in a department when AI is announced. It’s legitimate, and slogans won’t address it. What works is explaining precisely what the tool does and doesn’t do, then demonstrating it on the agents’ actual work.
- The model proposes; the agent decides
- No document produced by the model is released without human review and signature. This rule appears in the usage charter, and each customized model's system prompt reiterates it at the end of the response. It protects the organization (accountability) and the agent (who remains the author).
- Train on real-world cases
- Two hours per service, using the week’s letters and case files, not generic examples. The goal is for each agent to leave with three situations where the tool saves them time, and two where they shouldn’t use it.
- Explain hallucinations
- A local model invents legal references, code provisions, and figures with the same confidence as a correct answer. Agents must know that every legal reference produced by the model must be verified on Légifrance, and that RAG reduces this risk for the local authority’s documents without eliminating it.
- Make usage visible
- An internal convention: include the note “written with the help of an assistant, reviewed by [agent]” in working documents, never in official acts. It makes the practice less intimidating and facilitates oversight.
- Involve employee representatives
- Introducing an AI tool affects how work is organized. Informing the territorial social committee (CST) before deployment prevents it from learning about it through rumors and turning the project into a conflict.
- Measure without monitoring
- Track the number of active users and qualitative feedback, not individual productivity. Conversation logs are for troubleshooting and compliance, not agent evaluation, and the charter should say so.
In organizations that have taken this approach, adoption is not a problem: staff readily adopt a tool that removes the tedious part of a letter. The area requiring attention is elsewhere: some staff rely too heavily on the model and stop proofreading. Cross-checking during the first few weeks and regularly reminding users about observed hallucinations are the best safeguards.
#Pitfalls and troubleshooting
- Slow or stuttering responses
- The model exceeds VRAM capacity, and part of it runs on the processor. Check the GPU/CPU split with ollama ps. Choose a smaller model, use more aggressive quantization, or reduce the context size requested by Open WebUI.
- RAG gives irrelevant answers
- Most often, the documents are poorly segmented (scanned PDFs, tables, complex layouts). Convert them to clean text before uploading, remove obsolete versions of regulations, and test each collection with ten questions whose answers you know.
- The model does not know the local-territory vocabulary
- It doesn't know what a CCAS, DETR, or DSP means in your context. Add the local authority's glossary to the system prompt or to a dedicated collection.
- An agent pasted a social services file into a conversation
- This isn't an escape from the community, but it is out-of-scope processing. Delete the conversation, remind users of the charter, and if necessary adjust model access rights by user group in Open WebUI.
- An update that breaks everything
- Pin the versions: image tag Open WebUI, Ollama version, exact model list. Test updates on a secondary machine or virtual machine before touching the agent server.
- The server is reachable from outside
- Regularly check that no port forwarding or firewall rule exposes 11434 or the Open WebUI port on the city hall’s public address. Thousands of open Ollama servers are continuously indexed on the internet.
#Go further
This guide establishes the framework and basic deployment. These site guides explain the components you will work with next:
- Deploy an AI chatbot for your team on the intranet
- Details on the Nginx reverse proxy, authentication, monitoring, and conversation backups for a Open WebUI multi-user deployment.
- Secure your Ollama server
- Check your exposure, add authentication, TLS, and sound network practices. Read this before opening access beyond the local network.
- Local LLMs and GDPR: private-data compliance
- The complete legal framework (GDPR, the European AI Act, CNIL recommendations) for documenting the impact assessment and entering it in the register.
- Local RAG with Ollama without coding
- To understand what happens when you add documents to a collection and improve the quality of the welcome page's answers.
- Choose your quantization (Q4, Q5, Q8, FP16)
- To choose between a larger Q4 model and a smaller Q8 model based on the VRAM of the server you purchase.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.