Docker Model Runner: run LLMs with Docker, without Ollama
Docker Model Runner is the model launcher built into Docker Desktop and Docker Engine. Enable the feature, then pull an OCI-packaged model from Docker Hub with docker model pull ai/qwen3.5, chat with docker model run, and connect your applications to an OpenAI-compatible API (port 12434 on the host, or model-runner.docker.internal from a container). Under the hood, it's llama.cpp — like Ollama — but controlled through the docker CLI, with no separate daemon to install.
If Docker is already at the heart of your stack, installing Ollama alongside it is redundant. Docker Model Runner brings direct LLM execution to Docker: models become OCI artifacts that you pull like images, the docker model CLI runs them, and an OpenAI-compatible API serves your applications. This guide shows how to enable it, pull a model from Docker Hub, call it over HTTP, and above all when Docker Model Runner makes sense compared with Ollama—without overselling the tool.
#Why Docker Model Runner?
Docker Model Runner is Docker’s answer to Ollama: a way to run LLMs locally without leaving the Docker ecosystem. You no longer manage a separate daemon or a separate model directory—models are distributed as OCI artifacts, pulled from a registry exactly like container images, and controlled through a new family of docker model commands.
Technically, the inference engine under Model Runner is the same as the one used by Ollama, LM Studio, or Jan: llama.cpp. The difference is therefore not raw speed, but integration. If your workstation or server already runs Docker, Model Runner avoids adding another tool and naturally connects your containers to a local model.
- OCI-format models
- An LLM is pulled, versioned, and pushed like a Docker image. Same registry, same authentication, same instincts.
- OpenAI-compatible API
- Endpoints /engines/v1/chat/completions, /completions, /models. Any OpenAI client can connect to them by changing the base URL.
- No additional tools
- No Ollama to install alongside it. The Docker CLI is enough, and inference integrates with Compose.
- GPU used directly
- The engine runs on the host (Apple Silicon through Metal, NVIDIA through CUDA), not in a container, so acceleration is preserved.
#Prerequisites
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
- Recent Docker Desktop
- Model Runner arrived with Docker Desktop 4.40 (macOS Apple Silicon), then expanded to Windows with NVIDIA GPU support. Update to the latest available version.
- Or Docker Engine on Linux
- On a Linux server without Docker Desktop, Model Runner is installed through the docker-model-plugin package (see below).
- VRAM or unified memory
- In Q4_K_M, budget ≈2 GB for a 3B, ≈5 GB for a 7B, ≈9 GB for a 14B, and ≈19 GB for a 32B. On Mac, unified memory serves as VRAM.
- A recommended GPU
- RTX 3060 12 GB to get started, RTX 4070/4080 in the mid-range, RTX 4090 24 GB or an M4 Pro Mac (24–48 GB unified) for large models. CPU-only works, but slowly.
#1. Enable Docker Model Runner
The feature isn't always enabled by default. In Docker Desktop, you'll find it in the settings; from the command line, a single instruction is enough.
- 01Via Docker DesktopOpen Settings → AI (or “Beta features” depending on the version), then check “Enable Docker Model Runner”. To call the API from the host, also enable “Enable host-side TCP support” and note the proposed port (12434 by default).
- 02Via the CLIOne command activates the service and, optionally, opens the TCP port on the host so you can access the API from your machine without going through a container.
- 03Checkdocker model status confirme que le runner tourne. docker model version affiche la version du plugin installé.
On a Linux server with Docker Engine (without Docker Desktop), Model Runner is added as a plugin. On a Debian/Ubuntu distribution:
#2. Pull a model in OCI format
Docker hosts a library of OCI-packaged models under Docker Hub’s ai/ namespace. Pull them exactly like images, using docker model pull. The tags encode the model size and quantization.
- ai/smollm2
- Very small model, ideal for validating the installation in a few seconds even without a GPU.
- ai/qwen3.5 · ai/gemma4 · ai/granite4.2
- The solid all-rounders of 2026; choose the size based on your VRAM (2B, 4B, 9B, 12B…). Qwen 3.5 9B (≈6.6 GB in Q4) is the right default for 8 GB, all under the Apache 2.0 license.
- Quantization tags
- A tag such as 9B-Q4_K_M specifies the size and compression. Q4_K_M is the recommended quality/memory compromise; Q5_K_M and Q8_0 are larger.
The OCI format means these models can live in any compatible registry: Docker Hub, as well as a private enterprise registry. You can therefore push an internal model just as you push an image, with the same authentication and access policies.
#3. Chat with docker model run
Just as docker run launches a container, docker model run launches a conversation. Without a message argument, it opens an interactive chat in the terminal; with a prompt, it responds once and returns control—perfect for scripting.
The first call to a model loads it into memory; subsequent calls reuse the loaded instance. The engine automatically unloads the model after a period of inactivity to free VRAM, without requiring you to manage a daemon.
#4. The OpenAI-compatible API
The real advantage of Docker Model Runner, as with Ollama, is its OpenAI-compatible API. Any tool designed for OpenAI's API works by simply changing the base URL. Two addresses exist, depending on where you call from.
- From the host (TCP)
- http://localhost:12434/engines/v1/… si vous avez activé le support TCP côté hôte (port 12434 par défaut).
- From a container
- http://model-runner.docker.internal/engines/v1/… — un nom DNS interne résolu automatiquement dans le réseau Docker.
A direct chat call with curl from the host, once the TCP port is enabled:
The model field must match a model pulled locally. With OpenAI’s Python SDK, you only need to redirect base_url; the API key can be any string, as Model Runner does not require it locally.
#5. Connect a container to the model
This is where the advantage of Docker integration becomes clear: a container running your application can call the model through the internal DNS without exposing a port on the host. From code running inside a container, the base URL becomes model-runner.docker.internal.
In practice, pass the URL through an environment variable so the same code works locally (port 12434) and in a container (internal DNS). A docker-compose.yml snippet that injects the endpoint into the application service:
#Docker Model Runner or Ollama, depending on your workflow
Both run llama.cpp and expose an OpenAI-compatible API. The choice depends on the ecosystem you work in, not raw performance.
- Choose Model Runner
- When Docker is already your foundation: you want your containers to communicate with the LLM over the Docker network, distribute models through a private OCI registry, and avoid installing and maintaining another tool.
- Choose Model Runner
- For a Compose stack where the model is one service among others, versioned and deployed with the same practices as your images.
- Stay on Ollama
- When you want the broadest and most up-to-date model catalog, an active community and extensive documentation, and a tool that works identically on Windows, macOS, and Linux without Docker Desktop.
- Stay on Ollama
- For simple office use, outside any containerized context: Ollama on its port 11434 remains the most direct option, with an ecosystem of interfaces (Open WebUI, LM Studio) already wired into it.
#Troubleshooting
- “docker: 'model' is not a docker command”
- The plugin is not installed, or Model Runner is not enabled. Update Docker Desktop, enable the feature, or install docker-model-plugin on Docker Engine.
- The API doesn't respond on localhost:12434
- TCP support on the host side is not enabled. Run docker desktop enable model-runner --tcp 12434 again, or check the option in Settings → AI.
- A container does not include the model
- From a container, use model-runner.docker.internal, not localhost: localhost points to the container itself, not the host.
- Slow loading or “out of memory”
- The model exceeds your VRAM. Pull a lighter variant (Q4_K_M tag instead of Q5/Q8) or a smaller size, and check that the GPU is detected correctly.
- The GPU is not being used (Windows)
- NVIDIA GPU support on Windows arrived after the initial release. Make sure you're using a version of Docker Desktop that supports it and that your drivers are up to date.
#Go further
Docker Model Runner is best appreciated when connected to the other local AI building blocks already covered on the site:
- Install Ollama: Windows, macOS, and Linux
- To compare it directly with the reference alternative and its 11434 port, and decide which one fits your workflow.
- Q4, Q5, Q8: which quantization should you choose
- To correctly read OCI model tags (7B-Q4_K_M, etc.) and balance quality, speed, and memory before pulling.
- llama-server: a local OpenAI API with llama.cpp
- To see the same engine exposed differently, with finer control over GPU offloading.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.