Open WebUI: ChatGPT-like interface for Ollama
Open WebUI Ollama is currently the most practical combination for having a ChatGPT-style interface on your own machine, with no cloud dependency. This pure self-hosted stack turns your Ollama server into a multi-user web application, with conversation history, model management, and document RAG. This guide covers installing Open WebUI Ollama via Docker, network configuration, compatible models based on your VRAM, advanced features (RAG, multi-user support, MCP), and the limitations to know before an internal production deployment.
What is Open WebUI, and why pair it with Ollama
Open WebUI (formerly Ollama WebUI) is an open-source front end written in SvelteKit + FastAPI, released under the BSD-3 license on github.com/open-webui/open-webui. It provides a conversational interface equivalent to ChatGPT, but connected to a local inference backend. Ollama serves as that backend: a daemon that loads quantized GGUF models and exposes an OpenAI-compatible HTTP API on port 11434.
Combining the two provides a local LLM interface complete:
- Frontend : Open WebUI on port 8080 (chat, history, system prompts, Python functions)
- Backend : Ollama on port 11434 (GPU memory management, quantization, streaming)
- Storage : SQLite or PostgreSQL for conversations and embeddings
Unlike LM Studio, which is single-user and desktop-only, Open WebUI supports accounts, groups, per-model permissions, and a RAG mode with an integrated ChromaDB vector database. It's the preferred option for sharing an inference server among multiple collaborators.
Installation via Docker: the recommended method
L'installation Docker WebUI is the path supported by the Open WebUI team. It avoids Python dependency conflicts and enables atomic updates.
Prerequisites :
- Docker 24+ and Docker Compose v2
- Ollama installed on the host or containerized (see ollama.com/download)
- GPU NVIDIA with drivers ≥ 535 and nvidia-container-toolkit, or CPU for small models
Minimum command (Ollama already installed on the host) :
docker run -d -p 3000:8080 \
--add-host=host.docker.internal:host-gateway \
-v open-webui:/app/backend/data \
--name open-webui \
--restart always \
ghcr.io/open-webui/open-webui:main
The interface is then accessible at http://localhost:3000. The first account created is automatically promoted to administrator.
All-in-one docker-compose stack (Ollama + WebUI + GPU NVIDIA) : see the reference file at docs.openwebui.com/getting-started/quick-start. The service ollama must mount a persistent volume /root/.ollama to keep downloaded models between restarts; otherwise, you'll re-download several hundred gigabytes every docker compose down.
For AMD ROCm, use the image ghcr.io/open-webui/open-webui:main coupled with ollama/ollama:rocm. Performance on RX 7900 XTX is approximately 70–80% of a RTX 4090 in Q4_K_M (estimated figure, model-dependent).
Choose the right model for your hardware
Open WebUI displays all models present in the local Ollama registry. VRAM remains the limiting factor. Here are three tiers that align with consumer and professional hardware.
A RTX 4090 (24 GB of VRAM) is suitable for 30–70B models in Q4 with partial CPU offload:
- Qwen 2.5 72B Instruct (72B, Qwen License) — Q4 VRAM ~42 GB, ctx 131072. Requires CPU offload or a second card.
- Llama 3.3 70B Instruct (70B, Llama 3.3 Community) — ~40 GB Q4 VRAM. Offloading is required here too.
- DeepSeek R1 Distill Llama 70B — good reasoning performance, with an MMLU score competitive with the best 70B models (to be confirmed depending on the version).
A workstation with 2× RTX 6000 Ada (96 GB total) makes it possible to target compact MoEs:
- gpt-oss 120B (117B, Apache 2.0, OpenAI) — Q4 VRAM ~70 GB, ctx 128000. OpenAI's open model, native Ollama integration.
- Mistral Small 4 (119B, Apache 2.0) — 256000 context, excellent French-language compromise.
- Qwen 3.5 122B-A10B — 10B active MoE, very low latency for its size.
An 8× H100 server (640 GB) or cluster opens the door to the frontiers:
- DeepSeek V3.2 (685B, MIT) — ~410 GB Q4 VRAM
- Kimi K2.6 (1000B, Modified MIT) — Q4 VRAM ~600 GB, ctx 256000
- DeepSeek V4 Pro 1.6T — frontier MoE, 1M-token context
For precise sizing based on your card, use the quelllm.fr configurator that combines available VRAM, desired quantization, and target context length.
Advanced features: RAG, functions, MCP
Open WebUI goes beyond simple chat. Three capabilities are worth configuring during installation.
Built-in RAG : drop PDFs, DOCX files, TXT files, or URLs into a “Collection.” Open WebUI splits them (default chunk size 1500, overlap 100) and generates embeddings via sentence-transformers/all-MiniLM-L6-v2 locally or via Ollama (nomic-embed-text), and stores them in ChromaDB. During inference, the relevant chunks are injected into the context. For large databases, switch to PostgreSQL + pgvector via the variable VECTOR_DB=pgvector.
Python functions (Pipelines) : Open WebUI lets you run arbitrary code for pre- or post-processing. Example: automatically route prompts containing “code” to Qwen3-Coder-Next 80B-A3B, and the others toward Mistral Medium 3.5 128B. The pipelines run in a separate container on port 9099, providing useful isolation for security.
MCP (Model Context Protocol) support : since version 0.6, Open WebUI consumes MCP servers. See the official spec for available servers (filesystem, GitHub, Postgres, etc.).
Detailed comparison of frontends on quelllm.fr/compare/open-webui-vs-lm-studio.
Observed performance and tokens per second
Throughput depends on the model/GPU combination. A few reference measurements with Ollama 0.5+ and Open WebUI 0.6+ in Q4_K_M:
- Llama 3.1 70B on 2× RTX 4090 (tensor parallelism Ollama experimental): approximately 18–22 tokens/s during generation (estimated)
- gpt-oss 120B on H100 80GB: 60-80 tokens/s (to be confirmed depending on offload)
- Qwen 3 235B-A22B on 4× A100 80GB: 45–55 tokens/s thanks to the MoE architecture, which activates only 22B
Open WebUI's WebSocket streaming adds imperceptible latency (<10 ms). The bottleneck remains the inference itself. To optimize, enable OLLAMA_FLASH_ATTENTION=1 et OLLAMA_KV_CACHE_TYPE=q8_0 in the Ollama container environment: 30–40% memory savings on the context (see github.com/ollama/ollama/blob/main/docs/faq.md).
For reasoning models such as DeepSeek R1 671B, plan for 3 to 10× more tokens generated per request (internal chain-of-thought), meaning response times of several minutes even on powerful infrastructure.
Security and multi-user deployment
In an enterprise, three precautions are essential:
- TLS reverse proxy : put Traefik or Caddy in front of Open WebUI. The default port 8080 has no encryption.
- SSO authentication : Open WebUI supports OIDC (Keycloak, Authentik) via environment variables
OAUTH_*. Disable open account creation withENABLE_SIGNUP=false. - Model-based RBAC : from the admin interface, restrict heavy models (>40 GB VRAM) to senior groups to prevent a user from saturating the queue.
Audit the logs regularly /app/backend/data/audit.log that track prompts and uploads.
FAQ
Q: Does Open WebUI work without Ollama?
Yes. Open WebUI accepts any OpenAI-compatible API: vLLM, llama.cpp server, LiteLLM, TGI. Configure the URL under Settings → Connections. Ollama remains the simplest way to get started because its CLI handles downloading, GGUF quantization, and memory automatically. For high-performance deployment with mature tensor parallelism, vLLM or SGLang are preferable.
Q: What is the minimum VRAM needed to get started?
With 8 GB of VRAM (RTX 3060, 4060), comfortably run 7-8B models such as lightweight distilled variants. For 12-16 GB, target 13-14B models. Rule of thumb: VRAM ≈ parameters × 0.6 in Q4_K_M, plus 1-2 GB for 8K context. See quelllm.fr/guide/vram-quantification for details by quantization (Q4, Q5, Q8, FP16).
Q: How do you install Open WebUI without Docker?
Via pip install open-webui puis open-webui serve. This method is documented but less isolated: conflicts with other Python environments are possible, and updates are more delicate. Reserve it for development machines. For a single-user workstation, LM Studio or Jan may be simpler; compare on quelllm.fr/compare/lm-studio-vs-jan.
Q: Can Open WebUI connect to multiple Ollama servers?
Yes. In Settings → Connections, add several Ollama URLs (e.g. http://gpu-01:11434, http://gpu-02:11434). Open WebUI aggregates the available models. Routing is manual (the user chooses the model); for automatic load balancing, put LiteLLM in the proxy layer. Useful for distributing Mixtral 8x22B Instruct on a node and Llama 3.1 405B Instruct on another one.
Q: Are conversations encrypted at rest?
No by default. The SQLite database /app/backend/data/webui.db stores messages in plaintext. For encryption at rest, use a Docker volume on an encrypted filesystem (LUKS, ZFS native encryption), or migrate to PostgreSQL with Transparent Data Encryption. Network traffic can be encrypted through the TLS reverse proxy mentioned above.
Q: Does Open WebUI support vision and images?
Yes, with multimodal models. Upload an image to the conversation: Open WebUI encodes it in base64 and sends it to Ollama if the model supports vision. Compatible models in the catalog: Qwen 2.5 VL 72B, Qwen 3 VL 235B-A22B, LLaVA-OneVision 72B et Molmo 72B.
Conclusion
Open WebUI Ollama forms the reference stack for serving a ChatGPT-like interface in self-hosting, from laptop to GPU cluster. Docker installation takes fifteen minutes, while RBAC and RAG configuration can be added in a few hours. Model selection remains the decisive factor: precisely match parameters, quantization, and available VRAM. To explore the 249 models indexed according to your hardware constraints, browse the quelllm.fr complete catalog or leave it configurator recommend the optimal model/quantization pair for your GPU.