Beginner 12 minInterfaces

Open WebUI with Ollama: guide complet

Ollama runs in your terminal, which is efficient but not comfortable for everyday use. This Open WebUI + Ollama tutorial installs a complete local chat interface in a few minutes, similar to ChatGPT: persistent history, Markdown, attachments, built-in RAG over your documents, and multi-account management. Everything runs in a Docker container, with no system dependencies.

By Mohamed Meguedmi·Update 2026-08-27·Tested on Windows, macOS, and Linux

#Why Open WebUI

Open WebUI (formerly Ollama WebUI) has become the go-to frontend for self-hosted LLMs. It’s an open-source web application (permissive license) that natively talks to Ollama, but also to any OpenAI-compatible endpoint — LM Studio, vLLM, llama.cpp server, or even an OpenAI key if you have one.

Familiar interface
Sidebar with history, central chat area, model selector at the top. Anyone who has already opened ChatGPT will find their way around in 30 seconds.
Built-in RAG
Drag a PDF, .docx, .md, or .txt into the conversation: Open WebUI chunks it, embeds it, and uses it as context. No RAG stack to build manually.
Native multi-user support
Local accounts, admin/user/pending roles, and manual registration approval. Perfect for a team or family.
100% offline once installed
The container, UI, and models run on your machine. No mandatory telemetry and no outbound calls if you block OpenAI/HuggingFace in the settings.
Extensible
Python pipelines (functions, filters, custom RAG), MCP tools, web search integration (SearXNG, Tavily), TTS/STT, image generation through ComfyUI or Automatic1111.
i
Open WebUI ≠ Ollama
Ollama is the inference engine (the daemon that loads the model and generates tokens). Open WebUI is the interface that communicates with Ollama through its HTTP API. You can have Ollama without Open WebUI; the reverse is more complicated—Open WebUI needs a backend that serves the models.

#Prerequisites

The Local AI Kit

Open WebUI responds, connected to Ollama. The Local AI Kit turns it into your private ChatGPT for the whole household: multiple accounts (ch. 6), questions asked of your documents (ch. 8), and the list of what truly stays local (ch. 13).

  • Lifetime online access
  • PDF + files
  • Lifetime updates
Ollama installed and working
The daemon must listen on http://localhost:11434. Check with curl http://localhost:11434/api/tags — you should get JSON (empty or containing your models).
Docker Desktop or Docker Engine
Windows/macOS: Docker Desktop. Linux: docker-ce through your distro’s package manager. Compose v2 is included.
2 GB of free RAM
Open WebUI itself uses little memory (200–400 MB). Most of the RAM/VRAM will be used by Ollama, which loads the models.
A Ollama model already downloaded
If the list is empty, run ollama pull qwen3.5:4b or ollama pull granite4.2:8b before starting—otherwise there will be nothing to select in the UI.
→
Not Ollama yet?
If Ollama is not installed, start with the installation guide for your OS. The Windows/macOS/Linux tutorial for Ollama takes 3 minutes. Come back here afterward.

#1. One-command Docker installation

The official image is published on GitHub Container Registry. A single command is enough to start Open WebUI and connect it automatically to your local Ollama.

Linux / macOS — local Ollama
docker run -d \
  -p 3000:8080 \
  --add-host=host.docker.internal:host-gateway \
  -v open-webui:/app/backend/data \
  --name open-webui \
  --restart always \
  ghcr.io/open-webui/open-webui:main
Windows PowerShell
docker run -d `
  -p 3000:8080 `
  --add-host=host.docker.internal:host-gateway `
  -v open-webui:/app/backend/data `
  --name open-webui `
  --restart always `
  ghcr.io/open-webui/open-webui:main

Let’s break down the flags. Each one serves a specific purpose:

-p 3000:8080
Open WebUI listens on port 8080 inside the container. We publish it on port 3000 of your machine. You can access it via http://localhost:3000.
--add-host=host.docker.internal:host-gateway
Essential on Linux: allows the container to reach Ollama, running outside Docker, via the host.docker.internal hostname. On Windows/macOS, Docker Desktop already configures it.
-v open-webui:/app/backend/data
A named volume that persists conversation history, user accounts, and indexed documents. Without it, everything disappears when the container restarts.
--restart always
The container restarts automatically when the machine boots. Open WebUI becomes a permanent service.
ghcr.io/open-webui/open-webui:main
The main tag is the latest stable version. To pin a version, use :v0.5.0 (or the current release). In production, don’t depend on main.
!
First image: ~1.5 GB
The first docker run downloads the complete image. Allow several minutes depending on your connection. Subsequent starts will be instant.

#2. First connection and admin account

Once the container has started, open your browser to the URL below.

Local interface
http://localhost:3000
  1. 01
    Create the admin account
    On first launch, Open WebUI asks you to create an account. The very first registered user automatically becomes an administrator. Email and password—all of it stays local, in the Docker volume.
  2. 02
    Check available models
    Your Ollama models should appear in the selector at the top of the screen. If the list is empty, the connection to Ollama is failing (see section 3 below).
  3. 03
    Start a test conversation
    Select a model and type a message. If the response arrives as a stream, everything is connected. Otherwise, open Settings > Admin Panel > Connections to troubleshoot.
→
Make sure to note the admin password
There is no graphical recovery procedure. If you lose it, you must either edit the SQLite database in the Docker volume or destroy everything and start over. Store it in your password manager.

#3. Connect Open WebUI to Ollama

In 95% of cases, the connection happens automatically through host.docker.internal. If it doesn’t, here’s how to force it manually.

Go to Settings (icon in the lower-left corner) > Admin Panel > Connections > Ollama API. Enter the URL:

URL Ollama from the container
http://host.docker.internal:11434

Click the test button (the refresh icon next to the field). A green indicator confirms the connection. Your model list reloads immediately.

!
Linux: Ollama must listen on 0.0.0.0
By default on Linux, the Ollama daemon listens only on 127.0.0.1, making it invisible from Docker. Edit the systemd service (sudo systemctl edit ollama) and add Environment="OLLAMA_HOST=0.0.0.0:11434" under [Service]. Then sudo systemctl restart ollama. Remember your firewall rules if the machine is exposed.

To verify from the CLI that Ollama is reachable from outside the container:

Connectivity test
# Depuis l'hôte
curl http://localhost:11434/api/tags

# Depuis le conteneur Open WebUI (Linux)
docker exec -it open-webui curl http://host.docker.internal:11434/api/tags

#4. RAG on your documents in 2 minutes

This is probably the feature that justifies the installation on its own. Open WebUI includes a complete RAG pipeline: text extraction (PDF, DOCX, MD, TXT, HTML, source code), chunking, embeddings, vector search, and context injection.

#Method 1: Attachments on the fly

In a conversation, click the paperclip (or type # to browse indexed documents). Select a file—it is ingested, chunked, and embedded within seconds. The model can now answer questions about its contents.

i
Default embedding model
Open WebUI uses sentence-transformers/all-MiniLM-L6-v2 by default. It is fast but centered on English. For French content, switch to BAAI/bge-m3 or intfloat/multilingual-e5-large in Settings > Documents > Embedding Model. First download: ~1-2 GB.

#Method 2: Knowledge (persistent base)

For recurring use—internal documentation, knowledge bases, project archives—create a Knowledge. Workspace > Knowledge > Create Knowledge. Give it a name (e.g., "Product Docs"), bulk-upload your documents, and associate it with a custom model via Workspace > Models.

Default chunking
1,000 characters with an overlap of 100. Adjustable in Settings > Documents. For dense technical text, reduce it to 500/50. For narrative text, stay at 1500/200.
Top K
Number of chunks returned to the model. Default: 4. Increase to 6-8 for cross-cutting questions, or lower to 2-3 if the model starts to ramble.
Hybrid search
Can be enabled on the same page. Combines lexical BM25 with vector similarity. Essential for queries containing exact technical terms (product references, proper names, codes).
→
Choose a model suited to RAG
Injected context can weigh in at 2 to 8k tokens. A small model with a short context window saturates quickly. For serious RAG, prefer Qwen 3.5 9B (256k context, ≈6.6 GB), Granite 4.2 8B (128k, highly token-efficient, ≈5.3 GB), or Gemma 4 12B Q4 if you have the VRAM (≈7.6 GB).

#5. Multi-user and authentication

Open WebUI manages three roles: admin (everything), user (chat + their own knowledge bases), pending (account created but awaiting approval). The system is designed so an admin controls who joins the instance.

  1. 01
    Enable controlled registration
    Admin Panel > Settings > General. Set Default User Role to 'pending'. Every new registration will require your manual approval in Admin Panel > Users.
  2. 02
    Create users
    Your colleagues go to http://votre-ip:3000 and create an account. You see the request in Admin Panel > Users and approve it with one click. They can then sign in.
  3. 03
    Restrict models by user
    Workspace > Models > select a model > Visibility. You can make a model public, private, or expose it only to certain users (useful for a sensitive fine-tuned model).
  4. 04
    Force HTTPS if the instance is exposed
    Open WebUI doesn't handle TLS itself. Put Caddy, Traefik, or nginx in front of the container. Without HTTPS, don't expose it outside your LAN — passwords are transmitted in plaintext.
!
OAuth / LDAP: possible but advanced
Open WebUI supports OAuth providers (Google, Microsoft, GitHub) and LDAP through environment variables (OAUTH_*, LDAP_*). This is useful in business settings but requires solid expertise. For personal use or a small team, local accounts are more than sufficient.

#Open WebUI vs Msty vs LobeChat

Three mature interfaces share the market in 2026. Here's how to distinguish them based on your profile.

Open WebUI
The most complete and extensible option. RAG, Python pipelines, multi-user support, MCP, web search. Requires Docker. Ideal if you want ONE interface for an entire team.
Msty
Native desktop app (Win/Mac/Linux), zero Docker, 1-click installation. Excellent UX for solo use. Built-in RAG as well. Less extensible than Open WebUI. Ideal for a developer or curious user who wants to try it quickly.
LobeChat
More of a "visual ChatGPT clone." Polished, with plugins and an agent marketplace. Excellent multi-provider support. Less advanced RAG. Ideal if you switch between Ollama locally and several APIs (OpenAI, Anthropic, Mistral cloud).
Quick verdict
Solo + personal machine: Msty. Team + dedicated server: Open WebUI. Power user who wants a polished multi-provider frontend: LobeChat.

#Troubleshooting

Empty model list
Open WebUI does not connect to Ollama. Check that: (1) ollama list shows models, (2) curl http://localhost:11434/api/tags responds, and (3) on Linux, OLLAMA_HOST=0.0.0.0:11434 is defined. Test the URL in Admin Panel > Connections.
Error 502 Bad Gateway
The container starts incorrectly. docker logs open-webui shows the cause. Common causes: a read-only mounted volume, port 3000 already in use, or a conflict with a previous instance (docker rm -f open-webui, then restart).
Slow tokens per second
The bottleneck is in Ollama, not Open WebUI. ollama ps should show 100% GPU. If it’s CPU or partial, the model is spilling out of VRAM—switch to a smaller quantization (Q4_K_M rather than Q5_K_M).
Documents not indexed
The first upload downloads the embedding model (1–2 GB), which may take some time. Check docker logs open-webui. Also verify that the file does not exceed the maximum size (adjustable under Settings > Documents > Max Upload File Size).
Update
docker pull ghcr.io/open-webui/open-webui:main puis docker stop open-webui && docker rm open-webui et relancez la commande run d'origine. Le volume open-webui:/app/backend/data préserve vos données.
Backup
docker run --rm -v open-webui:/data -v $(pwd):/backup alpine tar czf /backup/openwebui-backup.tar.gz -C /data . crée une archive de tout votre historique, comptes, knowledges. À faire avant chaque update majeure.

#Go further

With Open WebUI installed and connected to Ollama, you have a complete local AI workstation. A few natural next steps:

Improve RAG
Open WebUI’s built-in RAG is quite good, but for large corpora or more specialized searches, the site's ChromaDB RAG guide shows how to build a dedicated pipeline that is more performant and tunable.
Choose the right quantization
Q4_K_M by default, but the tradeoff varies depending on your VRAM. The Q4/Q5/Q8 quantization guide details the rough figures—often decisive when choosing between running a 14B and staying with a 7B.
Compare with other frontends
If you are still undecided between Open WebUI, LibreChat, AnythingLLM, and SillyTavern, the comparative chat frontend guide lists their respective strengths on one page.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.