Beginner 14 minDocker

Ollama with Docker: installation and first model

Direct response

To run Ollama in Docker, launch the official ollama/ollama image with a volume for the models and port 11434: docker run -d -v ollama:/root/.ollama -p 11434:11434 --name ollama ollama/ollama. With an NVIDIA card, add --gpus=all after installing the NVIDIA Container Toolkit. On macOS, Docker Desktop provides no GPU access. Publish the port on 127.0.0.1 if the API must not be exposed to the network.

Ollama in a container is an isolated service that starts, updates, and removes with a single command, without touching the system. This guide starts with the official command, adds the NVIDIA GPU, an initial model, and a Compose file, then covers what tutorials leave out: who can actually reach port 11434, which version to pin, and why a GPU can disappear along the way.

By Mohamed Meguedmi·Update 2026-09-29·Tested on Windows, macOS, and Linux

#Ollama in Docker: what it really changes

The official image is called ollama/ollama and is hosted on Docker Hub, where it has surpassed 100 million downloads. It is based on Ubuntu 24.04 and directly starts the Ollama server: the OLLAMA_HOST variable is set to 0.0.0.0:11434, so the API listens on port 11434 inside the container. Four decisions remain. Publish this port to the host, mount a volume at /root/.ollama to preserve the models, grant GPU access with --gpus=all if you have a NVIDIA card, and choose the image version. The container starts with no models at all: you download them afterward with the ollama command, executed inside the container. These four decisions are reflected identically in a Compose file.

You gain an isolated installation, with no system service, described by a single file and easy to run alongside other containers (web interface, vector database, n8n). You pay with GPU access that must be configured, no GPU access under macOS, and a network port whose exposure you must control—a point most tutorials omit.

#Which command to use for your hardware

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

A single image covers every case; only the launch options change with the GPU. The table draws on Ollama's documentation, including the constraint that most often causes problems.

Docker run options by machine (documentation Ollama, September 2026)
MachineImage and optionsHost-side requirementsA limitation to know
Without a GPUollama/ollama, no optionsDocker onlyReserve this path for small models
NVIDIA, Linuxollama/ollama with --gpus=allDriver 550 or later, NVIDIA Container ToolkitCompute capability 5.0 to 6.2 cards: driver 570 minimum
NVIDIA, WindowsSame command, Docker DesktopWSL2 backend, WSL2-compatible driver, up-to-date WSL2 kernelWithout WSL2, no GPU access
AMD Radeon, Linuxollama/ollama:rocm with --device /dev/kfd --device /dev/driAMD ROCm v7 driverUnlisted card: HSA_OVERRIDE_GFX_VERSION, experimental
Other GPUs (Vulkan)ollama/ollama with --device /dev/kfd --device /dev/driNothing: Vulkan is included in the imageCan be disabled with OLLAMA_VULKAN=0
NVIDIA Jetson--gpus=all and JETSON_JETPACK=5 or 6JetPack 5 or 6Ollama does not infer the version
Mac (Docker Desktop)ollama/ollama, CPU onlyNoneNo GPU passthrough: prefer native installation

#Prerequisites

Docker
Docker Engine on Linux, Docker Desktop on Windows or macOS, with the docker compose command for section 5.
Memory
The model's size, plus the context, plus some headroom for the system. qwen3.5:9b weighs 6.6 GB: target 16 GB of RAM on a processor, not 8.
Disk
20 GB free for the image and one or two models.
GPU NVIDIA, optional
The driver is installed on the host, never in the container; the NVIDIA Container Toolkit provides the bridge.

#The five-step procedure

  1. 01
    Install Docker
    Docker Engine on Linux, Docker Desktop on Windows or macOS. The docker version command must return a response from both the client and the engine.
  2. 02
    Prepare the GPU NVIDIA (optional)
    Install the NVIDIA Container Toolkit, run nvidia-ctk runtime configure --runtime=docker, restart Docker, then test with docker run --rm --gpus all ubuntu nvidia-smi.
  3. 03
    Start the container
    docker run avec -d, --name ollama, -v ollama:/root/.ollama, -p 127.0.0.1:11434:11434 et, avec NVIDIA, --gpus=all.
  4. 04
    Download a model
    docker exec -it ollama ollama pull, suivi de qwen3.5:4b pour une carte de 8 Go, ou de qwen3.5:9b avec plus de marge.
  5. 05
    Check
    docker exec ollama ollama ps doit afficher 100% GPU dans la colonne PROCESSOR.

#1. The docker run command

The Ollama documentation gives this for the processor alone: docker run -d -v ollama:/root/.ollama -p 11434:11434 --name ollama ollama/ollama. The version below limits port exposure to the local machine, as the next section explains, and adds automatic restart.

Ollama Docker, processor version
docker run -d \
  --name ollama \
  -p 127.0.0.1:11434:11434 \
  -v ollama:/root/.ollama \
  --restart unless-stopped \
  ollama/ollama
-p 127.0.0.1:11434:11434
Publishes the container's port 11434 on the host's port, for the host machine itself only.
-v ollama:/root/.ollama
Named volume mounted where Ollama stores its models. It survives container deletion.
--restart unless-stopped
Restarts the container after Docker or the machine restarts, unless manually stopped.
Verification
docker ps --filter name=ollama
curl http://localhost:11434/api/version

The response is a JSON object with the version number, for example {"version":"0.34.4"}, the stable version as of September 29, 2026.

#Port 11434: who can really access it

When installed natively, Ollama listens on 127.0.0.1 by default: only the machine itself can talk to it. In the container, the image instead sets OLLAMA_HOST to 0.0.0.0:11434; otherwise, the port would be unreachable from outside. Protection therefore depends on how you publish the port.

According to Docker's documentation, publishing a container's port is unsafe by default: it becomes accessible to the outside world, not just the host. The command -p 11434:11434 binds it to all addresses on the machine. But Ollama's local API requires no authentication: anyone who can reach the port can list your models, download them, or occupy your GPU.

!
On Ubuntu, ufw doesn't protect you here
Docker diverts traffic from published ports before it reaches ufw rules. A rule that rejects port 11434 is ignored; Docker’s documentation describes this as bypassing firewall configuration. Publishing on 127.0.0.1 is the simplest workaround.
-p 11434:11434
All of the host's interfaces: reachable from the local network, or even from the Internet.
-p 127.0.0.1:11434:11434
Local machine only. This is the default choice for this guide.
-p 192.168.1.10:11434:11434
A single host address; replace it with yours.

To intentionally expose Ollama to a network, put an authenticated reverse proxy in front of it, as explained in the hardening guide cited at the end of the page.

#Change the port on the host side

Port 11434 is already in use if another instance is running, often the native application. Don’t touch the container side: change only the number on the left. With -p 127.0.0.1:11435:11434, your machine talks to Ollama on port 11435; inside the container, nothing changes.

#Connect to Ollama from another container

Two services from the same Compose file share a network where each joins by its service name. Open WebUI therefore uses http://ollama:11434 without publishing the port externally. If Ollama runs on the host instead, the Open WebUI documentation points out the pitfall: it listens on 127.0.0.1 and remains invisible from the container. You must reach it through host.docker.internal and make it listen elsewhere.

#2. Enable the GPU NVIDIA

Docker does not see your GPU by default. On the host, you need an up-to-date NVIDIA driver, the NVIDIA Container Toolkit, and then the --gpus flag. Ollama requires driver 550 or newer, and 570 for older cards with compute capability 5.0 to 6.2. Containers use the host’s driver; they do not bundle one.

NVIDIA Container Toolkit (Ubuntu, Debian)
curl -fsSL https://nvidia.github.io/libnvidia-container/gpgkey \
  | sudo gpg --dearmor -o /usr/share/keyrings/nvidia-container-toolkit-keyring.gpg
curl -fsSL https://nvidia.github.io/libnvidia-container/stable/deb/nvidia-container-toolkit.list \
  | sed 's#deb https://#deb [signed-by=/usr/share/keyrings/nvidia-container-toolkit-keyring.gpg] https://#g' \
  | sudo tee /etc/apt/sources.list.d/nvidia-container-toolkit.list
sudo apt-get update
sudo apt-get install -y nvidia-container-toolkit
sudo nvidia-ctk runtime configure --runtime=docker
sudo systemctl restart docker

The nvidia-ctk command modifies /etc/docker/daemon.json so Docker knows about the NVIDIA runtime, which is why Docker must be restarted. For Fedora or RHEL, the Ollama documentation provides the yum or dnf variant. Before touching Ollama, isolate the problem with a throwaway container: if it fails, Ollama will not see your GPU either.

GPU passthrough test
docker run --rm --gpus all ubuntu nvidia-smi

Then restart Ollama with GPU access. The volume is preserved: models that were already downloaded remain there.

Ollama Docker, GPU version NVIDIA
docker rm -f ollama

docker run -d \
  --name ollama \
  --gpus=all \
  -p 127.0.0.1:11434:11434 \
  -v ollama:/root/.ollama \
  --restart unless-stopped \
  ollama/ollama
!
Windows: the WSL2 backend is required
According to the Docker documentation, GPU access in Docker Desktop exists only on Windows with the WSL2 backend. You need a NVIDIA driver that supports WSL2, an up-to-date Windows installation, and the latest WSL2 kernel, obtained with wsl --update. To compare with Ollama installed on Windows, see the WSL2 or native guide.

For a Radeon on Linux, the rocm tag and the /dev/kfd and /dev/dri devices replace --gpus; the ROCm v7 driver still needs to be installed on the host.

Ollama Docker, AMD GPU (ROCm)
docker run -d \
  --device /dev/kfd --device /dev/dri \
  -v ollama:/root/.ollama \
  -p 127.0.0.1:11434:11434 \
  --name ollama \
  ollama/ollama:rocm

#3. First model

The container is running, but it's empty. The ollama command lives inside the container, so we call it with docker exec. This guide uses qwen3.5:9b, which the Ollama library describes as 6.6 GB, with a 256K context window and support for text and images.

Download a model
docker exec -it ollama ollama pull qwen3.5:9b

These 6.6 GB must fit in memory before the context even comes into play. On an 8 GB card, the margin is slim: the 4B variant is the prudent choice.

Weights of the qwen3.5 variants in the Ollama library (September 2026)
TagDisk footprintReference point for choosing
qwen3.5:4b3.4 GB8 GB card, with headroom for the context
qwen3.5:9b6.6 GB12 GB or more, or 16 GB of RAM on the processor
qwen3.5:27b17 GB24 GB card, moderate context
qwen3.5:35b24 GBDoes not fit on a 24 GB card with context: spills into RAM

The markers in the last column are rough estimates, not measurements: the actual space depends on context and quantization. To chat with the model, launch the container with ollama run dans.

Conversation
docker exec -it ollama ollama run qwen3.5:9b

Routine usage goes through the HTTP API: any client talks to the container like a native Ollama, and the documentation specifies that the API accepts a subset of the OpenAI format.

API call from the host
curl http://localhost:11434/api/generate -d '{
  "model": "qwen3.5:9b",
  "prompt": "Explique la quantification Q4 en deux phrases.",
  "stream": false
}'
Is the model on the GPU?
docker exec ollama ollama ps

In the PROCESSOR column, 100% GPU means the model is entirely on the card, 100% CPU means it is in system memory, and 48%/52% CPU/GPU means it is split between the two. Splitting it slows generation significantly: a smaller model is better than one that spills over.

#The context is configured on the container

The default context window depends on memory: 4k with less than 24 GiB of VRAM, 32k between 24 and 48 GiB, and 256k above that. The documentation recommends at least 64,000 tokens for agents and coding tools. A larger context uses more memory: pass -e OLLAMA_CONTEXT_LENGTH=8192 to docker run, then check ollama ps's CONTEXT column.

#4. Persistent volume: where your models live

The -v ollama:/root/.ollama creates a Docker volume named ollama, separate from the container. Docker's documentation confirms this: a volume persists after the container is deleted, allowing you to replace the image without downloading the models again. The docker volume inspect ollama command shows its Mountpoint, the location of the files on the host.

On Linux, the native installation stores its models in /usr/share/ollama/.ollama/models. This directory has no connection to the Docker volume: a model downloaded in one place does not appear in the other.

→
Store models on another drive
Replace the named volume with a host directory: -v /mnt/ssd/ollama:/root/.ollama. The files are directly readable, but you manage the directory permissions yourself.
Volume backup
docker run --rm \
  -v ollama:/data \
  -v $(pwd):/backup \
  alpine tar czf /backup/ollama-models.tar.gz -C /data .

#5. Docker Compose: the same thing, in a file

The file below does the same thing as the previous commands, can be reviewed at a glance, and can be versioned in git. It pins the image version: replace 0.34.4 with the latest stable version when you read this.

compose.yaml
services:
  ollama:
    image: ollama/ollama:0.34.4
    container_name: ollama
    ports:
      - "127.0.0.1:11434:11434"
    volumes:
      - ollama:/root/.ollama
    restart: unless-stopped
    deploy:
      resources:
        reservations:
          devices:
            - driver: nvidia
              count: all
              capabilities: [gpu]

volumes:
  ollama:

The deploy block reserves the GPU; according to the Compose documentation, the capabilities field is mandatory, otherwise deployment fails. Without NVIDIA GPU support, remove the entire block: the rest runs on the CPU.

Lifecycle
docker compose up -d        # démarrer
docker compose logs -f      # suivre les logs
docker compose pull         # télécharger l'image du tag indiqué
docker compose down         # arrêter, volume conservé
docker compose down -v      # arrêter et supprimer le volume : efface tous les modèles
i
Watch out for the -v option
According to the Docker documentation, docker compose down -v removes the named volumes declared in the file. The ollama volume is one of them: your models disappear.

#Update Ollama without losing your models

As of September 29, 2026, Ollama's versions page on GitHub lists 0.34.4 as the latest stable version, while 0.35.0 appears there as a preview. Yet the 0.35.0 tag already exists on Docker Hub, among rc tags for release candidate. Pin a stable version number instead of following the latest published tag: the update arrives when you decide.

With Compose, change the number in the file, then run docker compose pull and docker compose up -d. With docker run, pull the new image, remove the old container, and run the same command again: as long as the volume is the same, the models follow. Finally, check the version with curl http://localhost:11434/api/version.

#Docker or native installation: the choice

Both methods produce the same server on the same port. The table is based on the documentation for Ollama; it does not compare speeds because no published, citable measurement is available.

Docker and native installation, criterion by criterion
CriterionOllama in DockerOllama installed natively
UpdateChange the tag, then run docker compose pullAutomatic on macOS and Windows; on Linux, rerun the installation script
ModelsDocker volume or mounted directory/usr/share/ollama/.ollama/models sous Linux
Logsdocker logs ollamajournalctl -u ollama on Linux with systemd
Network exposureChoosing -p (127.0.0.1 or all addresses)127.0.0.1 by default, configurable with OLLAMA_HOST
GPU on MacNo accessDirect installation on the machine

Choose Docker to stack multiple services, pin a version, or share a workstation. Choose the native install on a Mac, or when Docker adds nothing for individual use.


#Troubleshooting

could not select device driver "nvidia"
The NVIDIA Container Toolkit is missing, or Docker was not restarted after nvidia-ctk runtime configure. Reconfigure it, restart Docker, and retest with docker run --rm --gpus all ubuntu nvidia-smi.
The GPU works, then Ollama falls back to the CPU
Documented symptom: the log reports GPU discovery failures after a while. Ollama recommends disabling systemd's cgroup management in Docker: add "exec-opts": ["native.cgroupdriver=cgroupfs"] to /etc/docker/daemon.json, then restart Docker.
GPU errors 3, 46, 100, or 999
Reload the UVM driver with sudo rmmod nvidia_uvm, then sudo modprobe nvidia_uvm, or restart the machine.
Port 11434 already in use
Another instance is listening, often the native application. Stop it, or publish on another port with -p 127.0.0.1:11435:11434.
Truncated responses
The default context depends on VRAM. Add -e OLLAMA_CONTEXT_LENGTH=8192 when launching, while monitoring memory.
Download blocked behind a proxy
Pass -e HTTPS_PROXY=https://proxy.example.com to the container. The documentation discourages HTTP_PROXY, which can disrupt clients.
Inaccessible from the local network
Normal with -p 127.0.0.1:11434:11434. Open the port only behind authentication.

#Frequently asked questions

FAQ
What is the official Docker image for Ollama?+
This is ollama/ollama on Docker Hub, published by Ollama. The latest tag identifies the current version, a numeric tag such as 0.34.4 pins a version, and the -rocm suffix selects the image for AMD GPUs. For long-term use, pin a stable numeric tag instead of latest so an update arrives only when you decide it should.
Can Ollama use the GPU in Docker on Windows or Mac?+
On Windows, yes for a NVIDIA card, with Docker Desktop using the WSL2 backend, a NVIDIA driver compatible with WSL2, and an up-to-date WSL2 kernel (wsl --update). On macOS, no: Docker Desktop offers neither GPU passthrough nor emulation, so the container runs on the processor. To use the Apple chip, install Ollama directly on the machine.
How do I update Ollama in Docker without losing the models?+
Models live in the volume, not in the container. With Compose, change the image tag, then run docker compose pull and docker compose up -d. With docker run, run docker pull, docker rm -f ollama, and rerun the same command with the same volume. Avoid docker compose down -v, which also deletes the named volumes from the file.
How do you connect Open WebUI to Ollama with Docker Compose?+
Place both services in the same Compose file: they share a network where each can be reached by name. In Open WebUI, set OLLAMA_BASE_URL to http://ollama:11434. If Ollama runs on the host instead, use host.docker.internal and make sure it is not listening only on 127.0.0.1. The Open WebUI guide details the configuration.
Is the Ollama API in a container protected by a password?+
No. The Ollama documentation states that the local API requires no authentication. With -p 11434:11434, Docker exposes the port on all host interfaces, and a ufw firewall does not filter it. Publish on 127.0.0.1, or put an authenticated reverse proxy in front of it, as described in the security guide.

#Go further

You have a functional, isolated Ollama with a model on the GPU. The logical next steps: a chat interface, security before any network sharing, then a production stack.

Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.