Intermediate 11 minDocker

Docker Model Runner: run LLMs with Docker, without Ollama

Direct response

Docker Model Runner is the model launcher built into Docker Desktop and Docker Engine. Enable the feature, then pull an OCI-packaged model from Docker Hub with docker model pull ai/qwen3.5, chat with docker model run, and connect your applications to an OpenAI-compatible API (port 12434 on the host, or model-runner.docker.internal from a container). Under the hood, it's llama.cpp — like Ollama — but controlled through the docker CLI, with no separate daemon to install.

If Docker is already at the heart of your stack, installing Ollama alongside it is redundant. Docker Model Runner brings direct LLM execution to Docker: models become OCI artifacts that you pull like images, the docker model CLI runs them, and an OpenAI-compatible API serves your applications. This guide shows how to enable it, pull a model from Docker Hub, call it over HTTP, and above all when Docker Model Runner makes sense compared with Ollama—without overselling the tool.

By Thomas P.·Update 2026-08-27·Tested on Windows, macOS, and Linux

#Why Docker Model Runner?

Docker Model Runner is Docker’s answer to Ollama: a way to run LLMs locally without leaving the Docker ecosystem. You no longer manage a separate daemon or a separate model directory—models are distributed as OCI artifacts, pulled from a registry exactly like container images, and controlled through a new family of docker model commands.

Technically, the inference engine under Model Runner is the same as the one used by Ollama, LM Studio, or Jan: llama.cpp. The difference is therefore not raw speed, but integration. If your workstation or server already runs Docker, Model Runner avoids adding another tool and naturally connects your containers to a local model.

i
In two words
Docker Model Runner = docker model pull/run/rm + an OpenAI-compatible API. Models are OCI artifacts hosted on Docker Hub (namespace ai/) or any compatible registry. The engine is llama.cpp, run on the host for direct GPU access—not in a container.
OCI-format models
An LLM is pulled, versioned, and pushed like a Docker image. Same registry, same authentication, same instincts.
OpenAI-compatible API
Endpoints /engines/v1/chat/completions, /completions, /models. Any OpenAI client can connect to them by changing the base URL.
No additional tools
No Ollama to install alongside it. The Docker CLI is enough, and inference integrates with Compose.
GPU used directly
The engine runs on the host (Apple Silicon through Metal, NVIDIA through CUDA), not in a container, so acceleration is preserved.

#Prerequisites

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates
Recent Docker Desktop
Model Runner arrived with Docker Desktop 4.40 (macOS Apple Silicon), then expanded to Windows with NVIDIA GPU support. Update to the latest available version.
Or Docker Engine on Linux
On a Linux server without Docker Desktop, Model Runner is installed through the docker-model-plugin package (see below).
VRAM or unified memory
In Q4_K_M, budget ≈2 GB for a 3B, ≈5 GB for a 7B, ≈9 GB for a 14B, and ≈19 GB for a 32B. On Mac, unified memory serves as VRAM.
A recommended GPU
RTX 3060 12 GB to get started, RTX 4070/4080 in the mid-range, RTX 4090 24 GB or an M4 Pro Mac (24–48 GB unified) for large models. CPU-only works, but slowly.
!
It isn't “the LLM in a container”
A common misconception: Model Runner doesn’t run the model inside a Docker container. The inference engine runs on the host, loaded on demand, to maintain direct GPU access. Docker orchestrates the download, OCI storage, and API, but inference remains native.

#1. Enable Docker Model Runner

The feature isn't always enabled by default. In Docker Desktop, you'll find it in the settings; from the command line, a single instruction is enough.

  1. 01
    Via Docker Desktop
    Open Settings → AI (or “Beta features” depending on the version), then check “Enable Docker Model Runner”. To call the API from the host, also enable “Enable host-side TCP support” and note the proposed port (12434 by default).
  2. 02
    Via the CLI
    One command activates the service and, optionally, opens the TCP port on the host so you can access the API from your machine without going through a container.
  3. 03
    Check
    docker model status confirme que le runner tourne. docker model version affiche la version du plugin installé.
Enable via CLI (Docker Desktop)
# Activer Model Runner
docker desktop enable model-runner

# Activer + exposer l'API sur le port hôte 12434
docker desktop enable model-runner --tcp 12434

# Vérifier l'état
docker model status

On a Linux server with Docker Engine (without Docker Desktop), Model Runner is added as a plugin. On a Debian/Ubuntu distribution:

Docker Engine — Linux
sudo apt-get update
sudo apt-get install docker-model-plugin

# Confirmer
docker model version
i
A new family of commands
docker model se comporte comme docker image ou docker container : pull, run, ls, rm, inspect, logs. Si vous connaissez la CLI Docker, vous connaissez déjà la logique de Model Runner.

#2. Pull a model in OCI format

Docker hosts a library of OCI-packaged models under Docker Hub’s ai/ namespace. Pull them exactly like images, using docker model pull. The tags encode the model size and quantization.

Pull models
# Un petit modèle pour tester rapidement
docker model pull ai/smollm2

# Un modèle plus capable, tag explicite
docker model pull ai/qwen3.5

# Une variante quantifiée précise (taille + quantization)
docker model pull ai/gemma4:12B-Q4_K_M

# Lister ce qui est stocké localement
docker model ls
ai/smollm2
Very small model, ideal for validating the installation in a few seconds even without a GPU.
ai/qwen3.5 · ai/gemma4 · ai/granite4.2
The solid all-rounders of 2026; choose the size based on your VRAM (2B, 4B, 9B, 12B…). Qwen 3.5 9B (≈6.6 GB in Q4) is the right default for 8 GB, all under the Apache 2.0 license.
Quantization tags
A tag such as 9B-Q4_K_M specifies the size and compression. Q4_K_M is the recommended quality/memory compromise; Q5_K_M and Q8_0 are larger.

The OCI format means these models can live in any compatible registry: Docker Hub, as well as a private enterprise registry. You can therefore push an internal model just as you push an image, with the same authentication and access policies.

→
Models from Hugging Face
Beyond the ai/ namespace, Model Runner can pull GGUFs directly from Hugging Face by prefixing the reference with hf.co/. Useful for a model that hasn't (yet) been published on Docker Hub.

#3. Chat with docker model run

Just as docker run launches a container, docker model run launches a conversation. Without a message argument, it opens an interactive chat in the terminal; with a prompt, it responds once and returns control—perfect for scripting.

Terminal
# Chat interactif
docker model run ai/qwen3.5

# Prompt unique (mode « one-shot », scriptable)
docker model run ai/qwen3.5 "Explique le format OCI en une phrase."

The first call to a model loads it into memory; subsequent calls reuse the loaded instance. The engine automatically unloads the model after a period of inactivity to free VRAM, without requiring you to manage a daemon.

Inspect and clean
# Détails d'un modèle (taille, quantization, architecture)
docker model inspect ai/qwen3.5

# Logs du moteur d'inférence
docker model logs

# Supprimer un modèle pour récupérer de l'espace disque
docker model rm ai/smollm2
i
On-demand loading
Model Runner does not keep all your models in VRAM. It loads the one you request, keeps it warm while in use, then releases it. This is similar to Ollama's behavior, without a resident process to monitor.

#4. The OpenAI-compatible API

The real advantage of Docker Model Runner, as with Ollama, is its OpenAI-compatible API. Any tool designed for OpenAI's API works by simply changing the base URL. Two addresses exist, depending on where you call from.

From the host (TCP)
http://localhost:12434/engines/v1/… si vous avez activé le support TCP côté hôte (port 12434 par défaut).
From a container
http://model-runner.docker.internal/engines/v1/… — un nom DNS interne résolu automatiquement dans le réseau Docker.

A direct chat call with curl from the host, once the TCP port is enabled:

Call /engines/v1/chat/completions
curl http://localhost:12434/engines/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "ai/qwen3.5",
    "messages": [
      {"role": "system", "content": "Tu réponds en français, de façon concise."},
      {"role": "user", "content": "Qu'\''est-ce qu'\''un artefact OCI ?"}
    ]
  }'

The model field must match a model pulled locally. With OpenAI’s Python SDK, you only need to redirect base_url; the API key can be any string, as Model Runner does not require it locally.

OpenAI Python client
from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:12434/engines/v1",
    api_key="docker",  # non vérifiée en local
)

resp = client.chat.completions.create(
    model="ai/qwen3.5",
    messages=[
        {"role": "user", "content": "Donne trois idées de noms pour un projet open source."}
    ],
)
print(resp.choices[0].message.content)
→
Migrate from Ollama
Ollama exposes the same family of endpoints on http://localhost:11434/v1. To switch an app from Ollama to Model Runner, change the base URL to http://localhost:12434/engines/v1 and change the model name. The rest of the OpenAI code stays unchanged.

#5. Connect a container to the model

This is where the advantage of Docker integration becomes clear: a container running your application can call the model through the internal DNS without exposing a port on the host. From code running inside a container, the base URL becomes model-runner.docker.internal.

From a container
curl http://model-runner.docker.internal/engines/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "ai/qwen3.5",
    "messages": [{"role": "user", "content": "Bonjour"}]
  }'

In practice, pass the URL through an environment variable so the same code works locally (port 12434) and in a container (internal DNS). A docker-compose.yml snippet that injects the endpoint into the application service:

docker-compose.yml
services:
  app:
    build: .
    environment:
      OPENAI_BASE_URL: http://model-runner.docker.internal/engines/v1
      OPENAI_API_KEY: docker
      MODEL_NAME: ai/qwen3.5
i
One model shared across services
Several containers can target the same endpoint, model-runner.docker.internal: the model is loaded once on the host and served to all of them. Ideal for a RAG stack or a multi-service back end that shares the same local LLM.

#Docker Model Runner or Ollama, depending on your workflow

Both run llama.cpp and expose an OpenAI-compatible API. The choice depends on the ecosystem you work in, not raw performance.

Choose Model Runner
When Docker is already your foundation: you want your containers to communicate with the LLM over the Docker network, distribute models through a private OCI registry, and avoid installing and maintaining another tool.
Choose Model Runner
For a Compose stack where the model is one service among others, versioned and deployed with the same practices as your images.
Stay on Ollama
When you want the broadest and most up-to-date model catalog, an active community and extensive documentation, and a tool that works identically on Windows, macOS, and Linux without Docker Desktop.
Stay on Ollama
For simple office use, outside any containerized context: Ollama on its port 11434 remains the most direct option, with an ecosystem of interfaces (Open WebUI, LM Studio) already wired into it.
i
Model Runner is young
Docker Model Runner is much newer than Ollama and evolving quickly: platform availability, commands, and the catalog change across Docker versions. Ollama remains, as of today, the most mature ecosystem. Check the official Docker documentation for the exact feature status.

#Troubleshooting

“docker: 'model' is not a docker command”
The plugin is not installed, or Model Runner is not enabled. Update Docker Desktop, enable the feature, or install docker-model-plugin on Docker Engine.
The API doesn't respond on localhost:12434
TCP support on the host side is not enabled. Run docker desktop enable model-runner --tcp 12434 again, or check the option in Settings → AI.
A container does not include the model
From a container, use model-runner.docker.internal, not localhost: localhost points to the container itself, not the host.
Slow loading or “out of memory”
The model exceeds your VRAM. Pull a lighter variant (Q4_K_M tag instead of Q5/Q8) or a smaller size, and check that the GPU is detected correctly.
The GPU is not being used (Windows)
NVIDIA GPU support on Windows arrived after the initial release. Make sure you're using a version of Docker Desktop that supports it and that your drivers are up to date.

#Go further

Docker Model Runner is best appreciated when connected to the other local AI building blocks already covered on the site:

Install Ollama: Windows, macOS, and Linux
To compare it directly with the reference alternative and its 11434 port, and decide which one fits your workflow.
Q4, Q5, Q8: which quantization should you choose
To correctly read OCI model tags (7B-Q4_K_M, etc.) and balance quality, speed, and memory before pulling.
llama-server: a local OpenAI API with llama.cpp
To see the same engine exposed differently, with finer control over GPU offloading.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.