Deploy an AI chatbot for your team on intranet
Your teams use ChatGPT without your permission, and every prompt could potentially send internal files to OpenAI. Deploying an enterprise AI chatbot on your intranet solves the problem in a few hours: a Ollama server, Open WebUI for multiple users, Nginx at the front with HTTPS, and a ChatGPT-like interface for the whole team, with no tokens leaving your network. This guide covers the complete stack, from hardware selection to conversation backups.
#Why deploy an enterprise AI chatbot on an intranet?
Three concrete reasons favor bringing it in-house: privacy (your prompts often contain code snippets, customer data, and financial information), cost (a ChatGPT Team subscription at €25/month/user quickly reaches €5,000/year for 20 people), and control (you choose the models, system prompts, and conversation logs).
A well-designed enterprise intranet AI chatbot stack fits on a single machine for teams of up to 30–50 people. Beyond that, separate the inference server from the frontend, but the architecture remains the same.
#Stack architecture
Deploy local AI at work: privacy, compliance, multi-user architecture, costs, the one-page memo for leadership.
- Lifetime online access
- PDF + files
- Lifetime updates
Four stacked components, each with a specific role:
- Ollama
- Inference daemon that hosts models (Qwen 3.5, Granite 4.2, Gemma 4, Mistral Small). Listens on http://localhost:11434 by default.
- Open WebUI
- ChatGPT-like multi-user frontend in Docker. Manages accounts, conversations, RAG, and admin/user roles.
- Nginx
- Reverse proxy at the front. Terminates HTTPS, applies rate limiting, and serves from a clean internal domain name.
- Watchtower + cron
- Automatic Docker image updates and daily backup of the Open WebUI volume (which contains accounts and conversations).
#Hardware and OS requirements
Sizing depends on the target model and concurrent load. Practical guidelines for a team:
- Small team (5–15 users, 8–9B model)
- 16 GB RAM, 8–12 GB GPU VRAM (RTX 3060 12GB, 4060 Ti 16GB). Qwen 3.5 9B Q4 (6.6 GB, 256k ctx) or Granite 4.2 8B (5.3 GB, very token-efficient).
- Medium-sized team (15–30 users, 20–24B model)
- 32 GB RAM, 16 GB VRAM GPU (RTX 4070 Ti Super, 5080). Mistral Small 24B Q4 (14 GB, good in French) or gpt-oss 20B (14 GB, very fast in MXFP4).
- Large team (30-50 users, 27-30B model)
- 64 GB RAM, a 24 GB VRAM GPU (RTX 3090/4090), or a 64 GB Mac Studio M5 Max. Qwen 3.8 27B Q4 (18 GB, 262k ctx, vision) or Granite 4.2 30B Q4 (18 GB, enterprise-focused).
- Beyond 50 simultaneous users
- Switch to vLLM or run multiple Ollama instances behind a load balancer. Beyond the scope of this guide.
On the OS side: Ubuntu Server 22.04 or 24.04 LTS remains the simplest choice. Docker Engine and NVIDIA drivers (with nvidia-container-toolkit if using a GPU) installed beforehand.
#1. Install Ollama on the server
Linux installation uses the official script. It detects the NVIDIA GPU and configures the systemd service automatically.
By default, the daemon listens only on localhost. For Docker (Open WebUI) to call it from its container, expose it on all local server interfaces. Edit the systemd override:
- OLLAMA_HOST=0.0.0.0:11434
- Listens on all interfaces. The OS firewall remains closed on 11434 from outside — only Open WebUI on the same machine can access it.
- OLLAMA_KEEP_ALIVE=30m
- Keeps the model loaded in VRAM for 30 minutes after the last prompt. Avoids costly reloads between user requests.
- OLLAMA_NUM_PARALLEL=4
- Number of parallel requests. 4 is a good starting point for an 8–9B model (such as Qwen 3.5 9B) on 12 GB of VRAM; drop to 2 on a smaller GPU.
#2. Open WebUI for multiple users with Docker
Open WebUI deploys with a single Docker command. The open-webui:/app/backend/data volume contains the entire database: accounts, conversations, prompts. That's what you'll need to back up.
- -p 127.0.0.1:3000:8080
- Open WebUI is reachable only from the server itself. Nginx will bridge to the internal network. Crucial for security.
- ENABLE_SIGNUP=false
- No one can create an account from the login page. You create users from the admin interface.
- DEFAULT_USER_ROLE=pending
- Every new account awaits admin approval. This prevents an employee from accidentally inviting an external user.
- OLLAMA_BASE_URL
- Points to Ollama running on the host. host.docker.internal resolves to the Docker gateway through --add-host.
The first account created through the http://serveur:3000 interface (via SSH forwarding or directly) automatically becomes an admin. Create it immediately after starting the container, before exposing the service.
#3. Nginx reverse proxy with internal SSL
To expose chat.entreprise.local cleanly, Nginx terminates HTTPS and forwards to Open WebUI at 127.0.0.1:3000. On an intranet, you can use a certificate issued by your internal PKI, or a self-signed certificate deployed to workstations through GPO.
- client_max_body_size 100M
- Essential for RAG: lets your users upload PDFs or documents up to 100 MB.
- Upgrade / Connection
- Enables WebSocket. Without these two lines, token-by-token streaming stops working and the interface appears frozen.
- proxy_read_timeout 600s
- A long generation request (summarizing a large document) can take several minutes. Nginx's default timeout (60s) cuts off the response midway.
#4. Authentication, roles, and onboarding
Open WebUI natively supports three roles: admin (configures everything), user (uses the chat), and pending (account created but not activated). For a small business, local authentication is sufficient. Beyond 30–40 people or for strict IT alignment, connect OIDC to your IdP (Keycloak, Authentik, Microsoft Entra).
- 01Create usersAdmin panel → Users → Add User. Enter the email, name, and initial password. The user will change it on their first login.
- 02Restrict models by roleIn Admin → Models, you can hide certain models from standard users. For example, keep Qwen 3.8 27B (the most VRAM-intensive) for admins only if you’ve loaded it.
- 03Enforce a default system promptAdmin → Settings → Interface → Default Prompt Suggestions. Ideal for guiding usage ('You answer in French, refuse sensitive topics, etc.').
- 04Disable external featuresAdmin → Settings → disable Web Search (otherwise Open WebUI calls DuckDuckGo) and Image Generation if you don't want any network output.
#5. Monitoring and conversation backups
Three things to monitor: server health (CPU/GPU/RAM), Ollama API health, and Open WebUI database integrity. And one thing to back up: the open-webui Docker volume.
#Quick monitoring
If you run Prometheus + Grafana, nvidia_smi_exporter for the GPU and node_exporter for the system cover 95% of cases. Open WebUI itself exposes /health over GET; add it to your uptime check.
#Daily backup of the Open WebUI volume
The open-webui Docker volume contains a SQLite database with all accounts, conversations, custom prompts, and documents indexed for RAG. Losing it means losing the team’s entire history.
The script keeps 14 days of history. For a proper disaster recovery plan, replicate /var/backups/open-webui to a NAS or encrypted S3 bucket every night using rsync or rclone.
#Troubleshooting
- Open WebUI sees no model at all
- The container cannot reach Ollama. Check docker exec open-webui curl http://host.docker.internal:11434/api/tags. If it times out, your Ollama is listening on 127.0.0.1—switch back to OLLAMA_HOST=0.0.0.0:11434.
- Choppy or stalled streaming
- Missing WebSocket headers in Nginx. Check that proxy_set_header Upgrade and Connection "upgrade" are present in the vhost config.
- 504 Gateway Timeout on long requests
- proxy_read_timeout is too low. Set it to at least 600s. For summaries of large documents, increase it to 1200s.
- GPU saturated, latency exploding
- Too many parallel requests. Lower OLLAMA_NUM_PARALLEL to 2 and increase OLLAMA_KEEP_ALIVE to avoid reloads.
- “no space left” error on /var/lib/docker
- The Ollama models aren't in Docker — the open-webui volume is what's growing. Run docker system prune, and monitor the RAG data (indexed documents quickly add up to several GB).
#Go further
You have an instance running for your team. Three natural paths to professionalize it:
- Document GDPR compliance
- The local LLM and GDPR guide covers auditing, the processing register, and CNIL recommendations—essential if your team handles personal data.
- Extend the tool with RAG
- The RAG guide with ChromaDB and Mistral shows how to connect your internal document database (wiki, exported SharePoint, contracts) to Open WebUI.
- Automate workflows
- The guide Automate with n8n and Ollama explains how to connect your chatbot to your business stack (incoming email, tickets, RSS) without any cloud leakage.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.