Intermediate 20 minDeployment

Deploy an AI chatbot for your team on intranet

Your teams use ChatGPT without your permission, and every prompt could potentially send internal files to OpenAI. Deploying an enterprise AI chatbot on your intranet solves the problem in a few hours: a Ollama server, Open WebUI for multiple users, Nginx at the front with HTTPS, and a ChatGPT-like interface for the whole team, with no tokens leaving your network. This guide covers the complete stack, from hardware selection to conversation backups.

By Mohamed Meguedmi·Update 2026-08-27·Tested on Windows, macOS, and Linux

#Why deploy an enterprise AI chatbot on an intranet?

Three concrete reasons favor bringing it in-house: privacy (your prompts often contain code snippets, customer data, and financial information), cost (a ChatGPT Team subscription at €25/month/user quickly reaches €5,000/year for 20 people), and control (you choose the models, system prompts, and conversation logs).

A well-designed enterprise intranet AI chatbot stack fits on a single machine for teams of up to 30–50 people. Beyond that, separate the inference server from the frontend, but the architecture remains the same.

i
What you get at the end
An internal HTTPS URL (https://chat.entreprise.local) accessible from every workstation on the intranet. Each employee has their own account, history, and folders. No data leaves the network. The whole setup takes less than 20 minutes once the server is prepared.

#Stack architecture

The AI at Work Kit

Deploy local AI at work: privacy, compliance, multi-user architecture, costs, the one-page memo for leadership.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

Four stacked components, each with a specific role:

Ollama
Inference daemon that hosts models (Qwen 3.5, Granite 4.2, Gemma 4, Mistral Small). Listens on http://localhost:11434 by default.
Open WebUI
ChatGPT-like multi-user frontend in Docker. Manages accounts, conversations, RAG, and admin/user roles.
Nginx
Reverse proxy at the front. Terminates HTTPS, applies rate limiting, and serves from a clean internal domain name.
Watchtower + cron
Automatic Docker image updates and daily backup of the Open WebUI volume (which contains accounts and conversations).
→
Why Open WebUI instead of LibreChat or AnythingLLM?
Open WebUI has the best native integration with Ollama (automatic model discovery, VRAM memory management), a robust multi-user mode with RBAC, and built-in RAG with no external dependency. Today, it is the default choice for an intranet deployment.

#Hardware and OS requirements

Sizing depends on the target model and concurrent load. Practical guidelines for a team:

Small team (5–15 users, 8–9B model)
16 GB RAM, 8–12 GB GPU VRAM (RTX 3060 12GB, 4060 Ti 16GB). Qwen 3.5 9B Q4 (6.6 GB, 256k ctx) or Granite 4.2 8B (5.3 GB, very token-efficient).
Medium-sized team (15–30 users, 20–24B model)
32 GB RAM, 16 GB VRAM GPU (RTX 4070 Ti Super, 5080). Mistral Small 24B Q4 (14 GB, good in French) or gpt-oss 20B (14 GB, very fast in MXFP4).
Large team (30-50 users, 27-30B model)
64 GB RAM, a 24 GB VRAM GPU (RTX 3090/4090), or a 64 GB Mac Studio M5 Max. Qwen 3.8 27B Q4 (18 GB, 262k ctx, vision) or Granite 4.2 30B Q4 (18 GB, enterprise-focused).
Beyond 50 simultaneous users
Switch to vLLM or run multiple Ollama instances behind a load balancer. Beyond the scope of this guide.

On the OS side: Ubuntu Server 22.04 or 24.04 LTS remains the simplest choice. Docker Engine and NVIDIA drivers (with nvidia-container-toolkit if using a GPU) installed beforehand.

!
Avoid the "server under the desk"
A team that depends on the tool cannot tolerate an outage caused by a power failure. Put the machine in the server room, on a UPS, with stable SSH access. If you have no infrastructure, a Mac mini M4 on a shelf makes an excellent quiet, power-efficient inference server (40 W at idle).

#1. Install Ollama on the server

Linux installation uses the official script. It detects the NVIDIA GPU and configures the systemd service automatically.

Ollama installation
curl -fsSL https://ollama.com/install.sh | sh

By default, the daemon listens only on localhost. For Docker (Open WebUI) to call it from its container, expose it on all local server interfaces. Edit the systemd override:

systemd configuration
sudo systemctl edit ollama
/etc/systemd/system/ollama.service.d/override.conf
[Service]
Environment="OLLAMA_HOST=0.0.0.0:11434"
Environment="OLLAMA_KEEP_ALIVE=30m"
Environment="OLLAMA_NUM_PARALLEL=4"
OLLAMA_HOST=0.0.0.0:11434
Listens on all interfaces. The OS firewall remains closed on 11434 from outside — only Open WebUI on the same machine can access it.
OLLAMA_KEEP_ALIVE=30m
Keeps the model loaded in VRAM for 30 minutes after the last prompt. Avoids costly reloads between user requests.
OLLAMA_NUM_PARALLEL=4
Number of parallel requests. 4 is a good starting point for an 8–9B model (such as Qwen 3.5 9B) on 12 GB of VRAM; drop to 2 on a smaller GPU.
Reload and download the model
sudo systemctl daemon-reload
sudo systemctl restart ollama
ollama pull qwen3.5:9b
→
Which model should a French team choose?
Qwen 3.5 9B (6.6 GB, 256k ctx, multimodal, Apache 2.0) offers the best quality/speed tradeoff in French for internal general users. For more demanding use cases (legal analysis, long-document summarization), move up to Mistral Small 24B Q4 (good in French) or Qwen 3.8 27B if your VRAM allows it.

#2. Open WebUI for multiple users with Docker

Open WebUI deploys with a single Docker command. The open-webui:/app/backend/data volume contains the entire database: accounts, conversations, prompts. That's what you'll need to back up.

Launch Open WebUI
docker run -d \
  --name open-webui \
  --restart always \
  -p 127.0.0.1:3000:8080 \
  -v open-webui:/app/backend/data \
  -e OLLAMA_BASE_URL=http://host.docker.internal:11434 \
  -e WEBUI_AUTH=true \
  -e ENABLE_SIGNUP=false \
  -e DEFAULT_USER_ROLE=pending \
  --add-host=host.docker.internal:host-gateway \
  ghcr.io/open-webui/open-webui:main
-p 127.0.0.1:3000:8080
Open WebUI is reachable only from the server itself. Nginx will bridge to the internal network. Crucial for security.
ENABLE_SIGNUP=false
No one can create an account from the login page. You create users from the admin interface.
DEFAULT_USER_ROLE=pending
Every new account awaits admin approval. This prevents an employee from accidentally inviting an external user.
OLLAMA_BASE_URL
Points to Ollama running on the host. host.docker.internal resolves to the Docker gateway through --add-host.

The first account created through the http://serveur:3000 interface (via SSH forwarding or directly) automatically becomes an admin. Create it immediately after starting the container, before exposing the service.

!
The first account = admin
If you leave Open WebUI open on the network before creating your admin account, the first person to sign up inherits full privileges. Always perform the initial setup locally or through an SSH tunnel, never directly exposed.

#3. Nginx reverse proxy with internal SSL

To expose chat.entreprise.local cleanly, Nginx terminates HTTPS and forwards to Open WebUI at 127.0.0.1:3000. On an intranet, you can use a certificate issued by your internal PKI, or a self-signed certificate deployed to workstations through GPO.

Install Nginx
sudo apt install nginx
/etc/nginx/sites-available/chat-intranet
server {
    listen 80;
    server_name chat.entreprise.local;
    return 301 https://$host$request_uri;
}

server {
    listen 443 ssl http2;
    server_name chat.entreprise.local;

    ssl_certificate     /etc/ssl/certs/chat-intranet.crt;
    ssl_certificate_key /etc/ssl/private/chat-intranet.key;
    ssl_protocols TLSv1.2 TLSv1.3;

    client_max_body_size 100M;

    location / {
        proxy_pass http://127.0.0.1:3000;
        proxy_http_version 1.1;
        proxy_set_header Host $host;
        proxy_set_header X-Real-IP $remote_addr;
        proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
        proxy_set_header X-Forwarded-Proto $scheme;

        proxy_set_header Upgrade $http_upgrade;
        proxy_set_header Connection "upgrade";

        proxy_read_timeout 600s;
        proxy_send_timeout 600s;
    }
}
client_max_body_size 100M
Essential for RAG: lets your users upload PDFs or documents up to 100 MB.
Upgrade / Connection
Enables WebSocket. Without these two lines, token-by-token streaming stops working and the interface appears frozen.
proxy_read_timeout 600s
A long generation request (summarizing a large document) can take several minutes. Nginx's default timeout (60s) cuts off the response midway.
Activate and reload
sudo ln -s /etc/nginx/sites-available/chat-intranet /etc/nginx/sites-enabled/
sudo nginx -t && sudo systemctl reload nginx

#4. Authentication, roles, and onboarding

Open WebUI natively supports three roles: admin (configures everything), user (uses the chat), and pending (account created but not activated). For a small business, local authentication is sufficient. Beyond 30–40 people or for strict IT alignment, connect OIDC to your IdP (Keycloak, Authentik, Microsoft Entra).

  1. 01
    Create users
    Admin panel → Users → Add User. Enter the email, name, and initial password. The user will change it on their first login.
  2. 02
    Restrict models by role
    In Admin → Models, you can hide certain models from standard users. For example, keep Qwen 3.8 27B (the most VRAM-intensive) for admins only if you’ve loaded it.
  3. 03
    Enforce a default system prompt
    Admin → Settings → Interface → Default Prompt Suggestions. Ideal for guiding usage ('You answer in French, refuse sensitive topics, etc.').
  4. 04
    Disable external features
    Admin → Settings → disable Web Search (otherwise Open WebUI calls DuckDuckGo) and Image Generation if you don't want any network output.
→
OIDC for further reading
Open WebUI supports OAUTH_PROVIDER_NAME, OAUTH_CLIENT_ID, OAUTH_CLIENT_SECRET, and OPENID_PROVIDER_URL as environment variables. In 10 minutes, your users can sign in with their Entra ID or Google Workspace account without managing local passwords.

#5. Monitoring and conversation backups

Three things to monitor: server health (CPU/GPU/RAM), Ollama API health, and Open WebUI database integrity. And one thing to back up: the open-webui Docker volume.

#Quick monitoring

Ollama health check
curl -s http://localhost:11434/api/tags | jq '.models | length'

If you run Prometheus + Grafana, nvidia_smi_exporter for the GPU and node_exporter for the system cover 95% of cases. Open WebUI itself exposes /health over GET; add it to your uptime check.

#Daily backup of the Open WebUI volume

The open-webui Docker volume contains a SQLite database with all accounts, conversations, custom prompts, and documents indexed for RAG. Losing it means losing the team’s entire history.

/usr/local/bin/backup-openwebui.sh
#!/bin/bash
set -e
BACKUP_DIR=/var/backups/open-webui
TIMESTAMP=$(date +%Y%m%d-%H%M)
mkdir -p "$BACKUP_DIR"

docker run --rm \
  -v open-webui:/data \
  -v "$BACKUP_DIR":/backup \
  alpine tar czf "/backup/openwebui-$TIMESTAMP.tar.gz" -C /data .

find "$BACKUP_DIR" -name 'openwebui-*.tar.gz' -mtime +14 -delete
Daily cron at 2 a.m.
sudo chmod +x /usr/local/bin/backup-openwebui.sh
echo "0 2 * * * root /usr/local/bin/backup-openwebui.sh" | sudo tee /etc/cron.d/openwebui-backup

The script keeps 14 days of history. For a proper disaster recovery plan, replicate /var/backups/open-webui to a NAS or encrypted S3 bucket every night using rsync or rclone.

!
Compliance and conversation logs
Conversations may contain personal or strategic data. Document the retention period in your GDPR register, and give users a way to purge their conversations (Open WebUI supports this natively). If possible, encrypt the Docker volume at the disk level (LUKS).

#Troubleshooting

Open WebUI sees no model at all
The container cannot reach Ollama. Check docker exec open-webui curl http://host.docker.internal:11434/api/tags. If it times out, your Ollama is listening on 127.0.0.1—switch back to OLLAMA_HOST=0.0.0.0:11434.
Choppy or stalled streaming
Missing WebSocket headers in Nginx. Check that proxy_set_header Upgrade and Connection "upgrade" are present in the vhost config.
504 Gateway Timeout on long requests
proxy_read_timeout is too low. Set it to at least 600s. For summaries of large documents, increase it to 1200s.
GPU saturated, latency exploding
Too many parallel requests. Lower OLLAMA_NUM_PARALLEL to 2 and increase OLLAMA_KEEP_ALIVE to avoid reloads.
“no space left” error on /var/lib/docker
The Ollama models aren't in Docker — the open-webui volume is what's growing. Run docker system prune, and monitor the RAG data (indexed documents quickly add up to several GB).

#Go further

You have an instance running for your team. Three natural paths to professionalize it:

Document GDPR compliance
The local LLM and GDPR guide covers auditing, the processing register, and CNIL recommendations—essential if your team handles personal data.
Extend the tool with RAG
The RAG guide with ChromaDB and Mistral shows how to connect your internal document database (wiki, exported SharePoint, contracts) to Open WebUI.
Automate workflows
The guide Automate with n8n and Ollama explains how to connect your chatbot to your business stack (incoming email, tickets, RSS) without any cloud leakage.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.