Ollama with Docker: installation and first model
To run Ollama in Docker, launch the official ollama/ollama image with a volume for the models and port 11434: docker run -d -v ollama:/root/.ollama -p 11434:11434 --name ollama ollama/ollama. With an NVIDIA card, add --gpus=all after installing the NVIDIA Container Toolkit. On macOS, Docker Desktop provides no GPU access. Publish the port on 127.0.0.1 if the API must not be exposed to the network.
Ollama in a container is an isolated service that starts, updates, and removes with a single command, without touching the system. This guide starts with the official command, adds the NVIDIA GPU, an initial model, and a Compose file, then covers what tutorials leave out: who can actually reach port 11434, which version to pin, and why a GPU can disappear along the way.
#Ollama in Docker: what it really changes
The official image is called ollama/ollama and is hosted on Docker Hub, where it has surpassed 100 million downloads. It is based on Ubuntu 24.04 and directly starts the Ollama server: the OLLAMA_HOST variable is set to 0.0.0.0:11434, so the API listens on port 11434 inside the container. Four decisions remain. Publish this port to the host, mount a volume at /root/.ollama to preserve the models, grant GPU access with --gpus=all if you have a NVIDIA card, and choose the image version. The container starts with no models at all: you download them afterward with the ollama command, executed inside the container. These four decisions are reflected identically in a Compose file.
You gain an isolated installation, with no system service, described by a single file and easy to run alongside other containers (web interface, vector database, n8n). You pay with GPU access that must be configured, no GPU access under macOS, and a network port whose exposure you must control—a point most tutorials omit.
#Which command to use for your hardware
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
A single image covers every case; only the launch options change with the GPU. The table draws on Ollama's documentation, including the constraint that most often causes problems.
| Machine | Image and options | Host-side requirements | A limitation to know |
|---|---|---|---|
| Without a GPU | ollama/ollama, no options | Docker only | Reserve this path for small models |
| NVIDIA, Linux | ollama/ollama with --gpus=all | Driver 550 or later, NVIDIA Container Toolkit | Compute capability 5.0 to 6.2 cards: driver 570 minimum |
| NVIDIA, Windows | Same command, Docker Desktop | WSL2 backend, WSL2-compatible driver, up-to-date WSL2 kernel | Without WSL2, no GPU access |
| AMD Radeon, Linux | ollama/ollama:rocm with --device /dev/kfd --device /dev/dri | AMD ROCm v7 driver | Unlisted card: HSA_OVERRIDE_GFX_VERSION, experimental |
| Other GPUs (Vulkan) | ollama/ollama with --device /dev/kfd --device /dev/dri | Nothing: Vulkan is included in the image | Can be disabled with OLLAMA_VULKAN=0 |
| NVIDIA Jetson | --gpus=all and JETSON_JETPACK=5 or 6 | JetPack 5 or 6 | Ollama does not infer the version |
| Mac (Docker Desktop) | ollama/ollama, CPU only | None | No GPU passthrough: prefer native installation |
#Prerequisites
- Docker
- Docker Engine on Linux, Docker Desktop on Windows or macOS, with the docker compose command for section 5.
- Memory
- The model's size, plus the context, plus some headroom for the system. qwen3.5:9b weighs 6.6 GB: target 16 GB of RAM on a processor, not 8.
- Disk
- 20 GB free for the image and one or two models.
- GPU NVIDIA, optional
- The driver is installed on the host, never in the container; the NVIDIA Container Toolkit provides the bridge.
#The five-step procedure
- 01Install DockerDocker Engine on Linux, Docker Desktop on Windows or macOS. The docker version command must return a response from both the client and the engine.
- 02Prepare the GPU NVIDIA (optional)Install the NVIDIA Container Toolkit, run nvidia-ctk runtime configure --runtime=docker, restart Docker, then test with docker run --rm --gpus all ubuntu nvidia-smi.
- 03Start the containerdocker run avec -d, --name ollama, -v ollama:/root/.ollama, -p 127.0.0.1:11434:11434 et, avec NVIDIA, --gpus=all.
- 04Download a modeldocker exec -it ollama ollama pull, suivi de qwen3.5:4b pour une carte de 8 Go, ou de qwen3.5:9b avec plus de marge.
- 05Checkdocker exec ollama ollama ps doit afficher 100% GPU dans la colonne PROCESSOR.
#1. The docker run command
The Ollama documentation gives this for the processor alone: docker run -d -v ollama:/root/.ollama -p 11434:11434 --name ollama ollama/ollama. The version below limits port exposure to the local machine, as the next section explains, and adds automatic restart.
- -p 127.0.0.1:11434:11434
- Publishes the container's port 11434 on the host's port, for the host machine itself only.
- -v ollama:/root/.ollama
- Named volume mounted where Ollama stores its models. It survives container deletion.
- --restart unless-stopped
- Restarts the container after Docker or the machine restarts, unless manually stopped.
The response is a JSON object with the version number, for example {"version":"0.34.4"}, the stable version as of September 29, 2026.
#Port 11434: who can really access it
When installed natively, Ollama listens on 127.0.0.1 by default: only the machine itself can talk to it. In the container, the image instead sets OLLAMA_HOST to 0.0.0.0:11434; otherwise, the port would be unreachable from outside. Protection therefore depends on how you publish the port.
According to Docker's documentation, publishing a container's port is unsafe by default: it becomes accessible to the outside world, not just the host. The command -p 11434:11434 binds it to all addresses on the machine. But Ollama's local API requires no authentication: anyone who can reach the port can list your models, download them, or occupy your GPU.
- -p 11434:11434
- All of the host's interfaces: reachable from the local network, or even from the Internet.
- -p 127.0.0.1:11434:11434
- Local machine only. This is the default choice for this guide.
- -p 192.168.1.10:11434:11434
- A single host address; replace it with yours.
To intentionally expose Ollama to a network, put an authenticated reverse proxy in front of it, as explained in the hardening guide cited at the end of the page.
#Change the port on the host side
Port 11434 is already in use if another instance is running, often the native application. Don’t touch the container side: change only the number on the left. With -p 127.0.0.1:11435:11434, your machine talks to Ollama on port 11435; inside the container, nothing changes.
#Connect to Ollama from another container
Two services from the same Compose file share a network where each joins by its service name. Open WebUI therefore uses http://ollama:11434 without publishing the port externally. If Ollama runs on the host instead, the Open WebUI documentation points out the pitfall: it listens on 127.0.0.1 and remains invisible from the container. You must reach it through host.docker.internal and make it listen elsewhere.
#2. Enable the GPU NVIDIA
Docker does not see your GPU by default. On the host, you need an up-to-date NVIDIA driver, the NVIDIA Container Toolkit, and then the --gpus flag. Ollama requires driver 550 or newer, and 570 for older cards with compute capability 5.0 to 6.2. Containers use the host’s driver; they do not bundle one.
The nvidia-ctk command modifies /etc/docker/daemon.json so Docker knows about the NVIDIA runtime, which is why Docker must be restarted. For Fedora or RHEL, the Ollama documentation provides the yum or dnf variant. Before touching Ollama, isolate the problem with a throwaway container: if it fails, Ollama will not see your GPU either.
Then restart Ollama with GPU access. The volume is preserved: models that were already downloaded remain there.
For a Radeon on Linux, the rocm tag and the /dev/kfd and /dev/dri devices replace --gpus; the ROCm v7 driver still needs to be installed on the host.
- Ollama and AMD GPUs: configuring ROCm step by step
- Ollama under WSL2 or native Windows: which should you choose?
#3. First model
The container is running, but it's empty. The ollama command lives inside the container, so we call it with docker exec. This guide uses qwen3.5:9b, which the Ollama library describes as 6.6 GB, with a 256K context window and support for text and images.
These 6.6 GB must fit in memory before the context even comes into play. On an 8 GB card, the margin is slim: the 4B variant is the prudent choice.
| Tag | Disk footprint | Reference point for choosing |
|---|---|---|
| qwen3.5:4b | 3.4 GB | 8 GB card, with headroom for the context |
| qwen3.5:9b | 6.6 GB | 12 GB or more, or 16 GB of RAM on the processor |
| qwen3.5:27b | 17 GB | 24 GB card, moderate context |
| qwen3.5:35b | 24 GB | Does not fit on a 24 GB card with context: spills into RAM |
The markers in the last column are rough estimates, not measurements: the actual space depends on context and quantization. To chat with the model, launch the container with ollama run dans.
Routine usage goes through the HTTP API: any client talks to the container like a native Ollama, and the documentation specifies that the API accepts a subset of the OpenAI format.
In the PROCESSOR column, 100% GPU means the model is entirely on the card, 100% CPU means it is in system memory, and 48%/52% CPU/GPU means it is split between the two. Splitting it slows generation significantly: a smaller model is better than one that spills over.
#The context is configured on the container
The default context window depends on memory: 4k with less than 24 GiB of VRAM, 32k between 24 and 48 GiB, and 256k above that. The documentation recommends at least 64,000 tokens for agents and coding tools. A larger context uses more memory: pass -e OLLAMA_CONTEXT_LENGTH=8192 to docker run, then check ollama ps's CONTEXT column.
#4. Persistent volume: where your models live
The -v ollama:/root/.ollama creates a Docker volume named ollama, separate from the container. Docker's documentation confirms this: a volume persists after the container is deleted, allowing you to replace the image without downloading the models again. The docker volume inspect ollama command shows its Mountpoint, the location of the files on the host.
On Linux, the native installation stores its models in /usr/share/ollama/.ollama/models. This directory has no connection to the Docker volume: a model downloaded in one place does not appear in the other.
#5. Docker Compose: the same thing, in a file
The file below does the same thing as the previous commands, can be reviewed at a glance, and can be versioned in git. It pins the image version: replace 0.34.4 with the latest stable version when you read this.
The deploy block reserves the GPU; according to the Compose documentation, the capabilities field is mandatory, otherwise deployment fails. Without NVIDIA GPU support, remove the entire block: the rest runs on the CPU.
#Update Ollama without losing your models
As of September 29, 2026, Ollama's versions page on GitHub lists 0.34.4 as the latest stable version, while 0.35.0 appears there as a preview. Yet the 0.35.0 tag already exists on Docker Hub, among rc tags for release candidate. Pin a stable version number instead of following the latest published tag: the update arrives when you decide.
With Compose, change the number in the file, then run docker compose pull and docker compose up -d. With docker run, pull the new image, remove the old container, and run the same command again: as long as the volume is the same, the models follow. Finally, check the version with curl http://localhost:11434/api/version.
#Docker or native installation: the choice
Both methods produce the same server on the same port. The table is based on the documentation for Ollama; it does not compare speeds because no published, citable measurement is available.
| Criterion | Ollama in Docker | Ollama installed natively |
|---|---|---|
| Update | Change the tag, then run docker compose pull | Automatic on macOS and Windows; on Linux, rerun the installation script |
| Models | Docker volume or mounted directory | /usr/share/ollama/.ollama/models sous Linux |
| Logs | docker logs ollama | journalctl -u ollama on Linux with systemd |
| Network exposure | Choosing -p (127.0.0.1 or all addresses) | 127.0.0.1 by default, configurable with OLLAMA_HOST |
| GPU on Mac | No access | Direct installation on the machine |
Choose Docker to stack multiple services, pin a version, or share a workstation. Choose the native install on a Mac, or when Docker adds nothing for individual use.
#Troubleshooting
- could not select device driver "nvidia"
- The NVIDIA Container Toolkit is missing, or Docker was not restarted after nvidia-ctk runtime configure. Reconfigure it, restart Docker, and retest with docker run --rm --gpus all ubuntu nvidia-smi.
- The GPU works, then Ollama falls back to the CPU
- Documented symptom: the log reports GPU discovery failures after a while. Ollama recommends disabling systemd's cgroup management in Docker: add "exec-opts": ["native.cgroupdriver=cgroupfs"] to /etc/docker/daemon.json, then restart Docker.
- GPU errors 3, 46, 100, or 999
- Reload the UVM driver with sudo rmmod nvidia_uvm, then sudo modprobe nvidia_uvm, or restart the machine.
- Port 11434 already in use
- Another instance is listening, often the native application. Stop it, or publish on another port with -p 127.0.0.1:11435:11434.
- Truncated responses
- The default context depends on VRAM. Add -e OLLAMA_CONTEXT_LENGTH=8192 when launching, while monitoring memory.
- Download blocked behind a proxy
- Pass -e HTTPS_PROXY=https://proxy.example.com to the container. The documentation discourages HTTP_PROXY, which can disrupt clients.
- Inaccessible from the local network
- Normal with -p 127.0.0.1:11434:11434. Open the port only behind authentication.
#Frequently asked questions
What is the official Docker image for Ollama?+
Can Ollama use the GPU in Docker on Windows or Mac?+
How do I update Ollama in Docker without losing the models?+
How do you connect Open WebUI to Ollama with Docker Compose?+
Is the Ollama API in a container protected by a password?+
#Go further
You have a functional, isolated Ollama with a model on the GPU. The logical next steps: a chat interface, security before any network sharing, then a production stack.
- Install Ollama on all systems: the general guide
- Troubleshoot Ollama: GPU not detected, slowdowns, out-of-memory errors
- Docker Model Runner: run LLMs with Docker, without Ollama
- Source: Ollama documentation, Docker page
- Source: ollama/ollama image on Docker Hub
- Source: Docker, port publishing
- Source: NVIDIA Container Toolkit, installation guide
- Source: Ollama documentation, supported GPU
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.