Advanced 22 minDocker

Deploy an LLM in production with Docker Compose

Running an LLM locally on your workstation takes ten minutes. Deploying an LLM in production with Docker Compose for a team—with an HTTPS URL, monitoring, backups, and real security—is a different undertaking. This guide assembles a complete stack (Ollama, Open WebUI, Qdrant, n8n, Traefik) in a single, commented docker-compose.yml, ready to deploy on a VPS or an on-premises server.

By Mohamed Meguedmi·Update 2026-08-27·Tested on Windows, macOS, and Linux

#Why this stack

The goal is to cover the four practical needs of an SMB or technical team adopting local AI: an inference engine (Ollama), a web user interface (Open WebUI), a vector database for RAG (Qdrant), and a no-code automation layer (n8n). We wrap everything with Traefik as a reverse proxy, which handles HTTPS through Let's Encrypt and routing.

Each service runs in its own isolated container, with its own persistent volumes. You can disable one without breaking the others, update a single component, or replicate the stack in a staging environment with one command.

i
What this guide is not
This is not a Kubernetes guide. For a team of 10–50 people on a single server (up to ~10 concurrent users on one GPU), Docker Compose is more than sufficient. Beyond that, or for multi-node high availability, look at k3s or Nomad.

#Target architecture

The Local Agents Kit

Agents that act on your machine: agentic Cline, MCP, n8n + Ollama, local automations.

  • Lifetime online access
  • PDF + files
  • Lifetime updates
Traefik (port 80/443)
Reverse proxy + automatic TLS termination via Let's Encrypt. All external traffic passes through it.
Open WebUI (internal)
Chat interface like ChatGPT, served at chat.votre-domaine.fr. Local authentication or OIDC.
Ollama (internal)
Inference daemon. NOT publicly exposed, accessible only from the Docker network.
Qdrant (internal)
Vector database for RAG. Storage for your documents' embeddings.
n8n
No-code automation, available at n8n.votre-domaine.fr.
Prometheus + Grafana
Optional, exposed internally for GPU/CPU/RAM monitoring.

#Prerequisites

A Linux server
Ubuntu 22.04 LTS or Debian 12, SSH root access, at least 16 GB of RAM, 200 GB of disk space.
One NVIDIA GPU recommended
RTX 3090 24 GB or RTX 4090 for comfortable 14B–32B models, or RTX 4070 12 GB for 7B. Without a GPU, stick to 3B models on the CPU.
A domain name
With DNS access to create subdomains. Required for Let's Encrypt.
Docker Engine + Compose v2
Installation via the official get.docker.com script. The Compose plugin has been included since 2022.
NVIDIA Container Toolkit
If you have a GPU, this is what allows Docker to expose it to containers.

#1. Prepare the server

Start with a fresh Ubuntu 22.04 installation. Three things to install: Docker, the NVIDIA Container Toolkit (if using a GPU), and a properly configured firewall.

Docker installation
curl -fsSL https://get.docker.com | sh
sudo usermod -aG docker $USER
newgrp docker
docker compose version
NVIDIA Container Toolkit
distribution=$(. /etc/os-release; echo $ID$VERSION_ID)
curl -fsSL https://nvidia.github.io/libnvidia-container/gpgkey | sudo gpg --dearmor -o /usr/share/keyrings/nvidia-container-toolkit-keyring.gpg
curl -s -L https://nvidia.github.io/libnvidia-container/$distribution/libnvidia-container.list | \
  sed 's#deb https://#deb [signed-by=/usr/share/keyrings/nvidia-container-toolkit-keyring.gpg] https://#g' | \
  sudo tee /etc/apt/sources.list.d/nvidia-container-toolkit.list
sudo apt-get update && sudo apt-get install -y nvidia-container-toolkit
sudo nvidia-ctk runtime configure --runtime=docker
sudo systemctl restart docker
Test whether Docker can see the GPU
docker run --rm --gpus all nvidia/cuda:12.4.0-base-ubuntu22.04 nvidia-smi
Minimal UFW firewall
sudo ufw default deny incoming
sudo ufw default allow outgoing
sudo ufw allow 22/tcp comment 'SSH'
sudo ufw allow 80/tcp comment 'HTTP - redir vers HTTPS'
sudo ufw allow 443/tcp comment 'HTTPS Traefik'
sudo ufw enable
!
Never expose Ollama directly
Ollama's port 11434 has no authentication. If you expose it to the Internet, anyone can use your GPU and exfiltrate your models through the API. Everything goes through Traefik with authentication.

#2. The commented docker-compose.yml

Create a working directory and the main file. This is the core of the deployment—each section is commented so you can adapt it without breaking everything.

Folder structure
mkdir -p /opt/llm-stack/{traefik,ollama,open-webui,qdrant,n8n}
cd /opt/llm-stack
touch docker-compose.yml .env
.env (to complete)
DOMAIN=votre-domaine.fr
ACME_EMAIL=admin@votre-domaine.fr
N8N_BASIC_AUTH_USER=admin
N8N_BASIC_AUTH_PASSWORD=changez-moi-vraiment
TRAEFIK_DASHBOARD_AUTH=admin:$$apr1$$xxxxxxx  # generé avec htpasswd
docker-compose.yml
name: llm-stack

networks:
  proxy:      # réseau public exposé par Traefik
  internal:   # réseau privé Ollama <-> WebUI <-> Qdrant

volumes:
  ollama_data:
  open_webui_data:
  qdrant_data:
  n8n_data:
  traefik_certs:

services:

  # --- Traefik : reverse proxy + Let's Encrypt automatique ---
  traefik:
    image: traefik:v3.2
    restart: unless-stopped
    command:
      - --providers.docker=true
      - --providers.docker.exposedbydefault=false
      - --entrypoints.web.address=:80
      - --entrypoints.web.http.redirections.entrypoint.to=websecure
      - --entrypoints.web.http.redirections.entrypoint.scheme=https
      - --entrypoints.websecure.address=:443
      - --certificatesresolvers.le.acme.email=${ACME_EMAIL}
      - --certificatesresolvers.le.acme.storage=/certs/acme.json
      - --certificatesresolvers.le.acme.tlschallenge=true
      - --api.dashboard=true
    ports:
      - 80:80
      - 443:443
    volumes:
      - /var/run/docker.sock:/var/run/docker.sock:ro
      - traefik_certs:/certs
    networks: [proxy]
    labels:
      - traefik.enable=true
      - traefik.http.routers.dashboard.rule=Host(`traefik.${DOMAIN}`)
      - traefik.http.routers.dashboard.service=api@internal
      - traefik.http.routers.dashboard.entrypoints=websecure
      - traefik.http.routers.dashboard.tls.certresolver=le
      - traefik.http.routers.dashboard.middlewares=dashboard-auth
      - traefik.http.middlewares.dashboard-auth.basicauth.users=${TRAEFIK_DASHBOARD_AUTH}

  # --- Ollama : moteur d'inférence, JAMAIS exposé en public ---
  ollama:
    image: ollama/ollama:latest
    restart: unless-stopped
    volumes:
      - ollama_data:/root/.ollama
    environment:
      - OLLAMA_HOST=0.0.0.0:11434
      - OLLAMA_KEEP_ALIVE=24h
      - OLLAMA_NUM_PARALLEL=2
    networks: [internal]
    deploy:
      resources:
        reservations:
          devices:
            - driver: nvidia
              count: all
              capabilities: [gpu]

  # --- Open WebUI : interface chat, exposée en HTTPS ---
  open-webui:
    image: ghcr.io/open-webui/open-webui:main
    restart: unless-stopped
    depends_on: [ollama]
    environment:
      - OLLAMA_BASE_URL=http://ollama:11434
      - WEBUI_AUTH=true
      - ENABLE_SIGNUP=false
    volumes:
      - open_webui_data:/app/backend/data
    networks: [proxy, internal]
    labels:
      - traefik.enable=true
      - traefik.http.routers.chat.rule=Host(`chat.${DOMAIN}`)
      - traefik.http.routers.chat.entrypoints=websecure
      - traefik.http.routers.chat.tls.certresolver=le
      - traefik.http.services.chat.loadbalancer.server.port=8080

  # --- Qdrant : base vectorielle pour le RAG ---
  qdrant:
    image: qdrant/qdrant:latest
    restart: unless-stopped
    volumes:
      - qdrant_data:/qdrant/storage
    networks: [internal]

  # --- n8n : automatisation no-code ---
  n8n:
    image: docker.n8n.io/n8nio/n8n:latest
    restart: unless-stopped
    environment:
      - N8N_BASIC_AUTH_ACTIVE=true
      - N8N_BASIC_AUTH_USER=${N8N_BASIC_AUTH_USER}
      - N8N_BASIC_AUTH_PASSWORD=${N8N_BASIC_AUTH_PASSWORD}
      - N8N_HOST=n8n.${DOMAIN}
      - N8N_PROTOCOL=https
      - WEBHOOK_URL=https://n8n.${DOMAIN}/
    volumes:
      - n8n_data:/home/node/.n8n
    networks: [proxy, internal]
    labels:
      - traefik.enable=true
      - traefik.http.routers.n8n.rule=Host(`n8n.${DOMAIN}`)
      - traefik.http.routers.n8n.entrypoints=websecure
      - traefik.http.routers.n8n.tls.certresolver=le
      - traefik.http.services.n8n.loadbalancer.server.port=5678
→
Why two networks
The internal network has no access to the reverse proxy. Ollama and Qdrant are the only services there—an attacker coming through Traefik has no way to reach their APIs directly. Open WebUI has one foot in both worlds: it receives public traffic through the proxy and communicates with Ollama over the internal network.

#3. Configure Traefik and DNS

Before the first docker compose up, create three A-record DNS entries pointing to your server's public IP:

chat.votre-domaine.fr
The user interface Open WebUI.
n8n.votre-domaine.fr
The n8n workflow editor.
traefik.votre-domaine.fr
The Traefik dashboard (protected by basic auth).

Generate the hash for the Traefik dashboard’s basic auth. Note: in .env, double the $ signs (docker-compose escaping).

Basic auth hash
sudo apt install apache2-utils
htpasswd -nb admin VotreMotDePasse | sed -e 's/\$/\$\$/g'
!
Let's Encrypt and rate limits
Let's Encrypt limits you to 50 certificates per domain per week. During testing, use --certificatesresolvers.le.acme.caserver=https://acme-staging-v02.api.letsencrypt.org/directory to point to the staging server. Change it once everything works.

#4. First launch and checks

  1. 01
    Start the stack
    From /opt/llm-stack, run docker compose up -d. Traefik will negotiate Let's Encrypt certificates within a few seconds — monitor the logs.
  2. 02
    Check service health
    docker compose ps doit afficher tous les services en running. Aucun ne doit redémarrer en boucle.
  3. 03
    Download a first model
    docker compose exec ollama ollama pull qwen3.5:9b (6,6 Go, 256k de contexte, le polyvalent 8 Go de référence en 2026). Pour un modèle plus costaud sur un GPU 24 Go : qwen3.8:27b (≈18 Go de VRAM, 262k de contexte, licence Apache 2.0).
  4. 04
    Test the interface
    Open https://chat.votre-domaine.fr. The first account created is the administrator. ENABLE_SIGNUP=false then blocks all public sign-ups.
  5. 05
    Verify that the GPU is being used
    docker compose exec ollama ollama ps : la colonne PROCESSOR doit afficher 100% GPU.
Live Traefik logs
docker compose logs -f traefik | grep -i acme

#5. Harden security

A misconfigured public stack gets scanned within the hour. Here are the six non-negotiable rules for a production LLM deployment with Docker Compose that won't end up in someone’s shodan.io.

Strong, unique passwords
Generate your secrets with openssl rand -base64 32. Never use the example password from .env.
Disable public sign-up
ENABLE_SIGNUP=false on Open WebUI. Without it, anyone can create an account and use your GPU.
Security headers via Traefik middleware
Add a secureHeaders middleware with HSTS, frame-deny, and content-type-nosniff. Three lines, huge impact.
Limit resources per container
Under deploy.resources.limits, set memory and cpus. Prevents a crashed service from bringing the entire server to its knees.
Containers in read-only mode whenever possible
Add read_only: true to Traefik and Qdrant. Use tmpfs for transient write directories.
Monthly updates
watchtower automates pulls, but prefer an explicit cron + git tag to stay in control of production versions.
Security-header middleware
# À ajouter sur le service open-webui en labels:
- traefik.http.routers.chat.middlewares=secure-headers
- traefik.http.middlewares.secure-headers.headers.stsSeconds=31536000
- traefik.http.middlewares.secure-headers.headers.stsIncludeSubdomains=true
- traefik.http.middlewares.secure-headers.headers.frameDeny=true
- traefik.http.middlewares.secure-headers.headers.contentTypeNosniff=true
- traefik.http.middlewares.secure-headers.headers.browserXssFilter=true
→
Fail2ban in front of Traefik
To block brute-force attacks on basic auth (Traefik dashboard, n8n), install fail2ban with a filter that parses Traefik logs. 5 attempts → 1 hour IP ban. Configuration in 10 lines.

#6. Prometheus + Grafana monitoring

Three metrics to monitor closely in production: the VRAM used by Ollama, the latency of Open WebUI requests, and the status of Let's Encrypt certificates (Traefik exposes everything at /metrics).

Block to add to docker-compose.yml
  prometheus:
    image: prom/prometheus:latest
    restart: unless-stopped
    volumes:
      - ./prometheus.yml:/etc/prometheus/prometheus.yml:ro
      - prometheus_data:/prometheus
    networks: [internal]

  nvidia-exporter:
    image: utkuozdemir/nvidia_gpu_exporter:latest
    restart: unless-stopped
    devices: [/dev/nvidiactl, /dev/nvidia0]
    volumes:
      - /usr/lib/x86_64-linux-gnu/libnvidia-ml.so.1:/usr/lib/x86_64-linux-gnu/libnvidia-ml.so.1:ro
      - /usr/bin/nvidia-smi:/usr/bin/nvidia-smi:ro
    networks: [internal]

  grafana:
    image: grafana/grafana:latest
    restart: unless-stopped
    volumes:
      - grafana_data:/var/lib/grafana
    networks: [proxy, internal]
    labels:
      - traefik.enable=true
      - traefik.http.routers.grafana.rule=Host(`grafana.${DOMAIN}`)
      - traefik.http.routers.grafana.entrypoints=websecure
      - traefik.http.routers.grafana.tls.certresolver=le
      - traefik.http.services.grafana.loadbalancer.server.port=3000
prometheus.yml
global:
  scrape_interval: 15s
scrape_configs:
  - job_name: 'traefik'
    static_configs:
      - targets: ['traefik:8080']
  - job_name: 'nvidia'
    static_configs:
      - targets: ['nvidia-exporter:9835']
  - job_name: 'qdrant'
    static_configs:
      - targets: ['qdrant:6333']

In Grafana, import dashboard ID 14574 (NVIDIA GPU Exporter) and 17346 (Traefik v3). You'll have a complete view in 5 minutes: VRAM, GPU temperatures, HTTP request rates by subdomain, and P95 latencies.

#7. Backups and updates

Five volumes contain all your critical data: ollama_data (models), open_webui_data (accounts + conversations + RAG), qdrant_data (vectors), n8n_data (workflows), traefik_certs (certificates).

/usr/local/bin/backup-llm-stack.sh
#!/bin/bash
set -e
BACKUP_DIR=/var/backups/llm-stack
TS=$(date +%Y%m%d-%H%M)
mkdir -p "$BACKUP_DIR"

for vol in open_webui_data qdrant_data n8n_data traefik_certs; do
  docker run --rm \
    -v llm-stack_${vol}:/data \
    -v "$BACKUP_DIR":/backup \
    alpine tar czf "/backup/${vol}-${TS}.tar.gz" -C /data .
done

find "$BACKUP_DIR" -name '*.tar.gz' -mtime +30 -delete
Daily cron at 3 a.m.
sudo chmod +x /usr/local/bin/backup-llm-stack.sh
echo "0 3 * * * root /usr/local/bin/backup-llm-stack.sh" | sudo tee /etc/cron.d/llm-stack-backup
i
Why ollama_data isn't backed up
Models quickly weigh tens of GB and can be re-downloaded with ollama pull. There’s no need to include them in daily backups — save the list instead: docker compose exec ollama ollama list > models.txt.

For updates, play it safe: pin versions (image: ollama/ollama:0.5.4 rather than :latest), snapshot the VPS before each update, and test on staging first.

Clean update
cd /opt/llm-stack
docker compose pull
docker compose up -d
docker image prune -f

#Troubleshooting

Traefik does not generate a certificate
Make sure ports 80 and 443 are open AND that DNS points to the server correctly. docker compose logs traefik | grep -i error reveals 90% of cases (DNS not propagated, rate limit, failed challenge).
Open WebUI displays 'No models found'
You haven't pulled a model into Ollama yet. docker compose exec ollama ollama pull qwen3.5:9b. Also verify OLLAMA_BASE_URL=http://ollama:11434 (the service name, not localhost).
GPU not detected in the container
nvidia-container-toolkit is not installed, or Docker was not restarted after configuration. Test with docker run --rm --gpus all nvidia/cuda:12.4.0-base-ubuntu22.04 nvidia-smi.
n8n redirects to HTTP in a loop
WEBHOOK_URL and N8N_PROTOCOL=https must be defined. Without them, n8n thinks it's using HTTP and the browser gets confused.
Grafana inaccessible behind Traefik
Add GF_SERVER_ROOT_URL=https://grafana.${DOMAIN} and GF_SERVER_SERVE_FROM_SUB_PATH=false to its environment.
Latency explodes with multiple users
OLLAMA_NUM_PARALLEL is too high for your VRAM. Lower it to 1 or 2 depending on your card. Monitor ollama ps while loading.

#Go further

Your stack is running, monitored, and backed up. Three natural directions for making it more professional:

Connect RAG to your documents
The RAG guide with ChromaDB and Mistral adapts easily to Qdrant—the same embedding and chunking logic, with a different vector database.
Automate business workflows
The Automate with n8n and Ollama guide describes ten concrete recipes (email translation, lead classification, RSS summarization) that directly leverage your stack.
Cover compliance
The local LLM and GDPR guide details the legal obligations (records, retention period, user rights) that must be documented for enterprise deployment.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.