Open WebUI: ChatGPT-like interface for Ollama

Open WebUI Ollama is currently the most practical combination for having a ChatGPT-style interface on your own machine, with no cloud dependency. This pure self-hosted stack turns your Ollama server into a multi-user web application, with conversation history, model management, and document RAG. This guide covers installing Open WebUI Ollama via Docker, network configuration, compatible models based on your VRAM, advanced features (RAG, multi-user support, MCP), and the limitations to know before an internal production deployment.

What is Open WebUI, and why pair it with Ollama

Open WebUI (formerly Ollama WebUI) is an open-source front end written in SvelteKit + FastAPI, released under the BSD-3 license on github.com/open-webui/open-webui. It provides a conversational interface equivalent to ChatGPT, but connected to a local inference backend. Ollama serves as that backend: a daemon that loads quantized GGUF models and exposes an OpenAI-compatible HTTP API on port 11434.

Combining the two provides a local LLM interface complete:

Unlike LM Studio, which is single-user and desktop-only, Open WebUI supports accounts, groups, per-model permissions, and a RAG mode with an integrated ChromaDB vector database. It's the preferred option for sharing an inference server among multiple collaborators.

Installation via Docker: the recommended method

L'installation Docker WebUI is the path supported by the Open WebUI team. It avoids Python dependency conflicts and enables atomic updates.

Prerequisites :

Minimum command (Ollama already installed on the host) :

docker run -d -p 3000:8080 \
  --add-host=host.docker.internal:host-gateway \
  -v open-webui:/app/backend/data \
  --name open-webui \
  --restart always \
  ghcr.io/open-webui/open-webui:main

The interface is then accessible at http://localhost:3000. The first account created is automatically promoted to administrator.

All-in-one docker-compose stack (Ollama + WebUI + GPU NVIDIA) : see the reference file at docs.openwebui.com/getting-started/quick-start. The service ollama must mount a persistent volume /root/.ollama to keep downloaded models between restarts; otherwise, you'll re-download several hundred gigabytes every docker compose down.

For AMD ROCm, use the image ghcr.io/open-webui/open-webui:main coupled with ollama/ollama:rocm. Performance on RX 7900 XTX is approximately 70–80% of a RTX 4090 in Q4_K_M (estimated figure, model-dependent).

Choose the right model for your hardware

Open WebUI displays all models present in the local Ollama registry. VRAM remains the limiting factor. Here are three tiers that align with consumer and professional hardware.

A RTX 4090 (24 GB of VRAM) is suitable for 30–70B models in Q4 with partial CPU offload:

A workstation with 2× RTX 6000 Ada (96 GB total) makes it possible to target compact MoEs:

An 8× H100 server (640 GB) or cluster opens the door to the frontiers:

For precise sizing based on your card, use the quelllm.fr configurator that combines available VRAM, desired quantization, and target context length.

Advanced features: RAG, functions, MCP

Open WebUI goes beyond simple chat. Three capabilities are worth configuring during installation.

Built-in RAG : drop PDFs, DOCX files, TXT files, or URLs into a “Collection.” Open WebUI splits them (default chunk size 1500, overlap 100) and generates embeddings via sentence-transformers/all-MiniLM-L6-v2 locally or via Ollama (nomic-embed-text), and stores them in ChromaDB. During inference, the relevant chunks are injected into the context. For large databases, switch to PostgreSQL + pgvector via the variable VECTOR_DB=pgvector.

Python functions (Pipelines) : Open WebUI lets you run arbitrary code for pre- or post-processing. Example: automatically route prompts containing “code” to Qwen3-Coder-Next 80B-A3B, and the others toward Mistral Medium 3.5 128B. The pipelines run in a separate container on port 9099, providing useful isolation for security.

MCP (Model Context Protocol) support : since version 0.6, Open WebUI consumes MCP servers. See the official spec for available servers (filesystem, GitHub, Postgres, etc.).

Detailed comparison of frontends on quelllm.fr/compare/open-webui-vs-lm-studio.

Observed performance and tokens per second

Throughput depends on the model/GPU combination. A few reference measurements with Ollama 0.5+ and Open WebUI 0.6+ in Q4_K_M:

Open WebUI's WebSocket streaming adds imperceptible latency (<10 ms). The bottleneck remains the inference itself. To optimize, enable OLLAMA_FLASH_ATTENTION=1 et OLLAMA_KV_CACHE_TYPE=q8_0 in the Ollama container environment: 30–40% memory savings on the context (see github.com/ollama/ollama/blob/main/docs/faq.md).

For reasoning models such as DeepSeek R1 671B, plan for 3 to 10× more tokens generated per request (internal chain-of-thought), meaning response times of several minutes even on powerful infrastructure.

Security and multi-user deployment

In an enterprise, three precautions are essential:

Audit the logs regularly /app/backend/data/audit.log that track prompts and uploads.

FAQ

Q: Does Open WebUI work without Ollama?

Yes. Open WebUI accepts any OpenAI-compatible API: vLLM, llama.cpp server, LiteLLM, TGI. Configure the URL under Settings → Connections. Ollama remains the simplest way to get started because its CLI handles downloading, GGUF quantization, and memory automatically. For high-performance deployment with mature tensor parallelism, vLLM or SGLang are preferable.

Q: What is the minimum VRAM needed to get started?

With 8 GB of VRAM (RTX 3060, 4060), comfortably run 7-8B models such as lightweight distilled variants. For 12-16 GB, target 13-14B models. Rule of thumb: VRAM ≈ parameters × 0.6 in Q4_K_M, plus 1-2 GB for 8K context. See quelllm.fr/guide/vram-quantification for details by quantization (Q4, Q5, Q8, FP16).

Q: How do you install Open WebUI without Docker?

Via pip install open-webui puis open-webui serve. This method is documented but less isolated: conflicts with other Python environments are possible, and updates are more delicate. Reserve it for development machines. For a single-user workstation, LM Studio or Jan may be simpler; compare on quelllm.fr/compare/lm-studio-vs-jan.

Q: Can Open WebUI connect to multiple Ollama servers?

Yes. In Settings → Connections, add several Ollama URLs (e.g. http://gpu-01:11434, http://gpu-02:11434). Open WebUI aggregates the available models. Routing is manual (the user chooses the model); for automatic load balancing, put LiteLLM in the proxy layer. Useful for distributing Mixtral 8x22B Instruct on a node and Llama 3.1 405B Instruct on another one.

Q: Are conversations encrypted at rest?

No by default. The SQLite database /app/backend/data/webui.db stores messages in plaintext. For encryption at rest, use a Docker volume on an encrypted filesystem (LUKS, ZFS native encryption), or migrate to PostgreSQL with Transparent Data Encryption. Network traffic can be encrypted through the TLS reverse proxy mentioned above.

Q: Does Open WebUI support vision and images?

Yes, with multimodal models. Upload an image to the conversation: Open WebUI encodes it in base64 and sends it to Ollama if the model supports vision. Compatible models in the catalog: Qwen 2.5 VL 72B, Qwen 3 VL 235B-A22B, LLaVA-OneVision 72B et Molmo 72B.

Conclusion

Open WebUI Ollama forms the reference stack for serving a ChatGPT-like interface in self-hosting, from laptop to GPU cluster. Docker installation takes fifteen minutes, while RBAC and RAG configuration can be added in a few hours. Model selection remains the decisive factor: precisely match parameters, quantization, and available VRAM. To explore the 249 models indexed according to your hardware constraints, browse the quelllm.fr complete catalog or leave it configurator recommend the optimal model/quantization pair for your GPU.

Article published on by Mohamed Meguedmi · Data source: /api/models.json · Content license: CC BY 4.0.

An error or update to report? Contribute.