Open WebUI with Ollama: guide complet
Ollama runs in your terminal, which is efficient but not comfortable for everyday use. This Open WebUI + Ollama tutorial installs a complete local chat interface in a few minutes, similar to ChatGPT: persistent history, Markdown, attachments, built-in RAG over your documents, and multi-account management. Everything runs in a Docker container, with no system dependencies.
#Why Open WebUI
Open WebUI (formerly Ollama WebUI) has become the go-to frontend for self-hosted LLMs. It’s an open-source web application (permissive license) that natively talks to Ollama, but also to any OpenAI-compatible endpoint — LM Studio, vLLM, llama.cpp server, or even an OpenAI key if you have one.
- Familiar interface
- Sidebar with history, central chat area, model selector at the top. Anyone who has already opened ChatGPT will find their way around in 30 seconds.
- Built-in RAG
- Drag a PDF, .docx, .md, or .txt into the conversation: Open WebUI chunks it, embeds it, and uses it as context. No RAG stack to build manually.
- Native multi-user support
- Local accounts, admin/user/pending roles, and manual registration approval. Perfect for a team or family.
- 100% offline once installed
- The container, UI, and models run on your machine. No mandatory telemetry and no outbound calls if you block OpenAI/HuggingFace in the settings.
- Extensible
- Python pipelines (functions, filters, custom RAG), MCP tools, web search integration (SearXNG, Tavily), TTS/STT, image generation through ComfyUI or Automatic1111.
#Prerequisites
Open WebUI responds, connected to Ollama. The Local AI Kit turns it into your private ChatGPT for the whole household: multiple accounts (ch. 6), questions asked of your documents (ch. 8), and the list of what truly stays local (ch. 13).
- Lifetime online access
- PDF + files
- Lifetime updates
- Ollama installed and working
- The daemon must listen on http://localhost:11434. Check with curl http://localhost:11434/api/tags — you should get JSON (empty or containing your models).
- Docker Desktop or Docker Engine
- Windows/macOS: Docker Desktop. Linux: docker-ce through your distro’s package manager. Compose v2 is included.
- 2 GB of free RAM
- Open WebUI itself uses little memory (200–400 MB). Most of the RAM/VRAM will be used by Ollama, which loads the models.
- A Ollama model already downloaded
- If the list is empty, run ollama pull qwen3.5:4b or ollama pull granite4.2:8b before starting—otherwise there will be nothing to select in the UI.
#1. One-command Docker installation
The official image is published on GitHub Container Registry. A single command is enough to start Open WebUI and connect it automatically to your local Ollama.
Let’s break down the flags. Each one serves a specific purpose:
- -p 3000:8080
- Open WebUI listens on port 8080 inside the container. We publish it on port 3000 of your machine. You can access it via http://localhost:3000.
- --add-host=host.docker.internal:host-gateway
- Essential on Linux: allows the container to reach Ollama, running outside Docker, via the host.docker.internal hostname. On Windows/macOS, Docker Desktop already configures it.
- -v open-webui:/app/backend/data
- A named volume that persists conversation history, user accounts, and indexed documents. Without it, everything disappears when the container restarts.
- --restart always
- The container restarts automatically when the machine boots. Open WebUI becomes a permanent service.
- ghcr.io/open-webui/open-webui:main
- The main tag is the latest stable version. To pin a version, use :v0.5.0 (or the current release). In production, don’t depend on main.
#2. First connection and admin account
Once the container has started, open your browser to the URL below.
- 01Create the admin accountOn first launch, Open WebUI asks you to create an account. The very first registered user automatically becomes an administrator. Email and password—all of it stays local, in the Docker volume.
- 02Check available modelsYour Ollama models should appear in the selector at the top of the screen. If the list is empty, the connection to Ollama is failing (see section 3 below).
- 03Start a test conversationSelect a model and type a message. If the response arrives as a stream, everything is connected. Otherwise, open Settings > Admin Panel > Connections to troubleshoot.
#3. Connect Open WebUI to Ollama
In 95% of cases, the connection happens automatically through host.docker.internal. If it doesn’t, here’s how to force it manually.
Go to Settings (icon in the lower-left corner) > Admin Panel > Connections > Ollama API. Enter the URL:
Click the test button (the refresh icon next to the field). A green indicator confirms the connection. Your model list reloads immediately.
To verify from the CLI that Ollama is reachable from outside the container:
#4. RAG on your documents in 2 minutes
This is probably the feature that justifies the installation on its own. Open WebUI includes a complete RAG pipeline: text extraction (PDF, DOCX, MD, TXT, HTML, source code), chunking, embeddings, vector search, and context injection.
#Method 1: Attachments on the fly
In a conversation, click the paperclip (or type # to browse indexed documents). Select a file—it is ingested, chunked, and embedded within seconds. The model can now answer questions about its contents.
#Method 2: Knowledge (persistent base)
For recurring use—internal documentation, knowledge bases, project archives—create a Knowledge. Workspace > Knowledge > Create Knowledge. Give it a name (e.g., "Product Docs"), bulk-upload your documents, and associate it with a custom model via Workspace > Models.
- Default chunking
- 1,000 characters with an overlap of 100. Adjustable in Settings > Documents. For dense technical text, reduce it to 500/50. For narrative text, stay at 1500/200.
- Top K
- Number of chunks returned to the model. Default: 4. Increase to 6-8 for cross-cutting questions, or lower to 2-3 if the model starts to ramble.
- Hybrid search
- Can be enabled on the same page. Combines lexical BM25 with vector similarity. Essential for queries containing exact technical terms (product references, proper names, codes).
#5. Multi-user and authentication
Open WebUI manages three roles: admin (everything), user (chat + their own knowledge bases), pending (account created but awaiting approval). The system is designed so an admin controls who joins the instance.
- 01Enable controlled registrationAdmin Panel > Settings > General. Set Default User Role to 'pending'. Every new registration will require your manual approval in Admin Panel > Users.
- 02Create usersYour colleagues go to http://votre-ip:3000 and create an account. You see the request in Admin Panel > Users and approve it with one click. They can then sign in.
- 03Restrict models by userWorkspace > Models > select a model > Visibility. You can make a model public, private, or expose it only to certain users (useful for a sensitive fine-tuned model).
- 04Force HTTPS if the instance is exposedOpen WebUI doesn't handle TLS itself. Put Caddy, Traefik, or nginx in front of the container. Without HTTPS, don't expose it outside your LAN — passwords are transmitted in plaintext.
#Open WebUI vs Msty vs LobeChat
Three mature interfaces share the market in 2026. Here's how to distinguish them based on your profile.
- Open WebUI
- The most complete and extensible option. RAG, Python pipelines, multi-user support, MCP, web search. Requires Docker. Ideal if you want ONE interface for an entire team.
- Msty
- Native desktop app (Win/Mac/Linux), zero Docker, 1-click installation. Excellent UX for solo use. Built-in RAG as well. Less extensible than Open WebUI. Ideal for a developer or curious user who wants to try it quickly.
- LobeChat
- More of a "visual ChatGPT clone." Polished, with plugins and an agent marketplace. Excellent multi-provider support. Less advanced RAG. Ideal if you switch between Ollama locally and several APIs (OpenAI, Anthropic, Mistral cloud).
- Quick verdict
- Solo + personal machine: Msty. Team + dedicated server: Open WebUI. Power user who wants a polished multi-provider frontend: LobeChat.
#Troubleshooting
- Empty model list
- Open WebUI does not connect to Ollama. Check that: (1) ollama list shows models, (2) curl http://localhost:11434/api/tags responds, and (3) on Linux, OLLAMA_HOST=0.0.0.0:11434 is defined. Test the URL in Admin Panel > Connections.
- Error 502 Bad Gateway
- The container starts incorrectly. docker logs open-webui shows the cause. Common causes: a read-only mounted volume, port 3000 already in use, or a conflict with a previous instance (docker rm -f open-webui, then restart).
- Slow tokens per second
- The bottleneck is in Ollama, not Open WebUI. ollama ps should show 100% GPU. If it’s CPU or partial, the model is spilling out of VRAM—switch to a smaller quantization (Q4_K_M rather than Q5_K_M).
- Documents not indexed
- The first upload downloads the embedding model (1–2 GB), which may take some time. Check docker logs open-webui. Also verify that the file does not exceed the maximum size (adjustable under Settings > Documents > Max Upload File Size).
- Update
- docker pull ghcr.io/open-webui/open-webui:main puis docker stop open-webui && docker rm open-webui et relancez la commande run d'origine. Le volume open-webui:/app/backend/data préserve vos données.
- Backup
- docker run --rm -v open-webui:/data -v $(pwd):/backup alpine tar czf /backup/openwebui-backup.tar.gz -C /data . crée une archive de tout votre historique, comptes, knowledges. À faire avant chaque update majeure.
#Go further
With Open WebUI installed and connected to Ollama, you have a complete local AI workstation. A few natural next steps:
- Improve RAG
- Open WebUI’s built-in RAG is quite good, but for large corpora or more specialized searches, the site's ChromaDB RAG guide shows how to build a dedicated pipeline that is more performant and tunable.
- Choose the right quantization
- Q4_K_M by default, but the tradeoff varies depending on your VRAM. The Q4/Q5/Q8 quantization guide details the rough figures—often decisive when choosing between running a 14B and staying with a 7B.
- Compare with other frontends
- If you are still undecided between Open WebUI, LibreChat, AnythingLLM, and SillyTavern, the comparative chat frontend guide lists their respective strengths on one page.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.