OpenClaw with Ollama: connecting a model local
This guide shows how to connect OpenClaw to Ollama to run the assistant on a local model: provider declaration, server address, context window requirements, and choosing a model that can call tools. Half the work is avoiding failures that display no error: truncated context, tools never called, or a model missing from the list. The commands follow the documentation for Ollama and OpenClaw; review them before pasting because both change quickly. This page contains no in-house testing, speed measurements, or model rankings.
#OpenClaw and Ollama: who does what
OpenClaw is a gateway: a process that receives your messages from a messaging service, sends them to a language model, and executes the actions requested by that model. Ollama is a model server: it loads a model into memory and responds on port 11434 of the machine, at http://localhost:11434 by default. Connecting one to the other means declaring Ollama as the provider in OpenClaw, then designating a local model as the agent's primary model.
According to the OpenClaw documentation, this connection uses the native API of Ollama (the /api/chat endpoint), which supports streaming responses and tool calling. This detail matters more than it might seem: we’ll see that a mistyped address is enough to switch the gateway to another mode, where tools no longer work.
- Ollama
- It loads the model, allocates a context window for it, and generates text. It determines how much memory is consumed.
- OpenClaw
- On every turn, it sends the system instruction, the description of available tools, and the conversation history, then executes the tools the model requests.
- The model
- Must keep a long instruction in mind and know how to request a tool in the expected format. Not all local models can do this.
- What doesn't change
- Messaging, assistant memory, token, and gateway security remain configured in OpenClaw, regardless of the model provider.
This integration is more delicate than a chat interface. A chat sends a few lines to the model; an agent immediately sends it several thousand tokens of instructions and tool definitions, before your first message. Ollama's default settings are designed for the first case, not the second.
#Prerequisites
Agents that act on your machine: agentic Cline, MCP, n8n + Ollama, local automations.
- Lifetime online access
- PDF + files
- Lifetime updates
- OpenClaw installed
- A gateway that starts up and passes diagnostics. Installation is not covered here: see our guide, "Installing OpenClaw with Docker."
- Ollama installed and up to date
- The ollama launch command used below exists only in recent versions. Installation is covered in our guide, “Install Ollama.”
- Memory for the model and its context
- Q4_K_M guidelines for weights alone: about 5 GB for a model with 7 billion parameters, 9 GB for 14 billion, and 19 GB for 32 billion. The context window requested by an agent is added to this figure.
- Terminal access
- On the machine hosting the gateway, and on the one hosting Ollama if it’s not the same machine.
The last command should return the list of installed models in JSON format. A refused connection means Ollama isn’t running: launch the application, or run ollama serve in a terminal. There’s no point going further until this response arrives.
#Step 1: Plan for 64,000 tokens of context
This is the setting that causes the most installations to fail, and it is configured on the Ollama side, not the OpenClaw side. The page that Ollama dedicates to OpenClaw says the assistant needs a large context window and recommends at least 64 000 tokens with a local model. Its page on context length gives the same value for agents, web search, and coding tools.
Or Ollama chooses its default context window based on available video memory: according to the same documentation, about 4,000 tokens under 24 GiB of VRAM, 32,000 between 24 and 48 GiB, and 256,000 starting at 48 GiB. On a 12- or 16-GB card, the server therefore starts with a window sixteen times smaller than recommended. Nothing indicates this: Ollama truncates anything beyond the limit without an error message.
If Ollama is already running as an application (macOS, Windows), do not run this command: a second server would conflict on port 11434. Set the context length in the application settings. On Linux, when Ollama has been installed as a systemd service, the variable is declared in the service itself.
#Step 2: choose a model that can call tools
An agent acts only through its tools: reading a file, running a command, searching the web. A model that cannot formulate a tool request will politely answer your messages, but will never do anything. This page does not rank models; it gives you the criteria to check before connecting one.
- The “tools” capability
- The ollama show affiche command has a Capabilities section. It must contain tools. In the Ollama library, the corresponding filter is available at https://ollama.com/search?c=tools.
- A sufficient native window
- The same command displays the model's maximum context length. A model designed for 8,000 or 32,000 tokens cannot follow the 64,000-token recommendation, regardless of the server setting.
- A realistic memory budget
- Weights and context must fit together in VRAM or in a Mac's unified memory. On a 12 GB card (RTX 3060, RTX 4070), this points to models significantly smaller than those the card accepts for simple chat; 16 GB (RTX 4080) and 24 GB (RTX 4090) leave more headroom.
- Long-term stability
- An agent chains together several tool calls per request. Very small models more often get the format or tool wrong. No fact sheet can replace testing with your own requests, starting with low-stakes ones.
The name gpt-oss:20b is used as an example throughout this guide: replace it with the model you selected. Ollama's integration page keeps an up-to-date list of models suggested for OpenClaw, which changes as new releases come out; it's better to refer to that than rely on a static list here.
#Step 3: configure the Ollama provider in OpenClaw
There are two approaches. The first is an Ollama command that writes the configuration for you. The second is to declare the provider yourself in the OpenClaw configuration; this is essential as soon as the gateway runs in Docker or Ollama is on another machine.
#Quick path: ollama launch openclaw
According to the Ollama documentation, this command lets you choose a model, configures OpenClaw to use Ollama, and starts the gateway; if it is already running, it reloads the new configuration automatically. The project's former name remains accepted: ollama launch clawdbot is an alias. The command is intended for an OpenClaw installation directly on the machine, with the openclaw command available in the terminal. It does not replace step 1: the same page asks you to record the server context.
#Manual route: declare the provider yourself
OpenClaw's documentation first describes an automatic discovery mode. We provide a dummy key, since Ollama doesn't require one, and OpenClaw queries the local instance at http://127.0.0.1:11434 to find the installed models.
A model is specified as ollama/ followed by the exact name displayed by ollama list, including the tag. Prefer putting it in the configuration rather than using an environment variable when the gateway runs as a service: a variable exported in your terminal is not passed to a process started by the system. With a Docker installation, prefix every openclaw command with docker compose run --rm openclaw-cli.
The second mode is explicit declaration in the ~/.openclaw/openclaw.json file, written in JSON5. It is used when Ollama runs somewhere other than the gateway machine, when a model does not appear in the list, or when you want to set the context window advertised to the agent yourself.
- baseUrl
- The address of server Ollama, including the port, with nothing after it. This is the only line to change when Ollama runs on another machine.
- api: "ollama"
- Explicitly request the native API of Ollama, the one that handles tool calls.
- apiKey
- A dummy value. It only serves to activate the provider.
- contextWindow
- The window announced to OpenClaw, which uses it to manage the history length. It must match what Ollama actually loads, not what the model would theoretically accept.
- maxTokens
- The maximum length of a response.
- cost
- Zero cost: a local model is not billed per token.
- agents.defaults.model.primary
- The model used by default by the agent, in the form ollama/model-name.
This example follows the structure provided in OpenClaw’s documentation; the contextWindow and maxTokens values are ours and should be adjusted for your model. Two things to remember. First, set reasoning to true for a reasoning model. Second, according to the same documentation, automatic discovery is disabled as soon as an explicit models.providers.ollama entry exists: every model you want to use must then appear in the models list.
#Gateway in Docker or Ollama on another machine
Inside a container, localhost refers to the container itself. An OpenClaw gateway launched with Docker therefore cannot see the host machine’s Ollama at http://localhost:11434: the connection is refused even though everything works from your terminal. The solution depends on where Ollama is running.
- Docker Desktop (macOS, Windows)
- The name host.docker.internal refers to the host machine from inside the container. Set http://host.docker.internal:11434 as the baseUrl in the explicit declaration.
- Docker Engine on Linux
- This name does not exist by default: you need to add it to the service with extra_hosts, as shown below. Ollama must also listen on an interface the container can reach, which is not the case with its original setting, limited to the loopback interface.
- Ollama on another machine
- Put this machine’s address on your local network or VPN in baseUrl, and likewise configure Ollama to listen on this machine.
The Compose file is an example from us, not an excerpt from the OpenClaw documentation: compare the service name with the docker-compose.yml for your version. The OLLAMA_HOST variable, however, is described in the FAQ for Ollama. Understand its implications: with 0.0.0.0, the server listens on all the machine's interfaces, and the Ollama API requires no authentication. A firewall must limit port 11434 to the Docker or local network, and this port must never be reachable from the Internet. Our guide to securing a Ollama server details these rules.
#Step 4: verify the end-to-end connection
An agent that responds “hello” proves nothing: that response requires neither a tool nor context. Useful verification proceeds layer by layer, from the model server to messaging.
- 01Test tool calling on Ollama aloneSend the server a question accompanied by a fictitious tool, using the command below. The response must contain a tool_calls field that names the tool and passes it an argument. If the model responds with a sentence, it is not suitable for an agent.
- 02Control what OpenClaw can seeThe command openclaw models list should display your model as ollama/model-name, and openclaw doctor should not report a provider error.
- 03Request an action, not an answerFrom the control interface or your messaging app, send a request that forces the agent to use a tool, such as listing the files in its workspace. It must actually do it, not describe what it would do.
- 04See what Ollama loadedImmediately after this exchange, run ollama ps on the server machine and read the CONTEXT and PROCESSOR columns.
In ollama ps output, the CONTEXT column shows the window actually allocated to the loaded model. If it displays 4096 when you intended 64,000, the setting from step 1 was not applied, regardless of what OpenClaw’s configuration says. The PROCESSOR column shows the split between GPU and CPU: 100% GPU is what you want; a mixed split indicates that the model and its context exceed video memory.
#Silent failures: symptoms and causes
Clear-cut errors (connection refused, model not found) appear in the logs. The failures below are more costly because the assistant keeps responding—it just responds incorrectly.
- The assistant ignores its instructions or responds off-topic
- Most likely cause: the context was truncated. The system prompt and tool definitions exceed the window loaded by Ollama, which cuts part of it off without warning. Check the CONTEXT column in ollama ps and resume step 1.
- JSON is displayed instead of the action
- The model correctly formulated a tool call, but the gateway received it as text. This indicates a /v1 endpoint or a provider configured in OpenAI-compatible mode. Switch back to the native endpoint and api: "ollama".
- It describes what it would do without doing anything
- The model does not declare the tools capability, or it is too limited to use it in the middle of a long instruction. Rerun the direct test on Ollama described in step 4; if it fails, change models.
- The model does not appear in openclaw models list
- Three possibilities. The provider isn't enabled (the dummy key is missing, or the variable wasn't passed to the service). An explicit models.providers.ollama entry exists and doesn't list this model. Or the model doesn't declare tool calling: according to the documentation we're familiar with, automatic discovery only includes models that declare it, a behavior that may have changed across versions.
- The context setting has no effect
- The OLLAMA_CONTEXT_LENGTH variable was exported in a terminal while Ollama was running as a service or application: the server never saw it. Declare it in the service or in the application settings, then restart Ollama.
- Responses take a very long time, or never arrive
- Either the model spills over onto the processor (the PROCESSOR column in ollama ps), or it was unloaded after a period of inactivity and reloads with every message: by default, Ollama keeps a model in memory for five minutes. The OLLAMA_KEEP_ALIVE variable extends this period.
- Everything works in the terminal, nothing through the gateway
- The gateway runs in a container and looks for Ollama on its own localhost. See the section on Docker.
- “Model context window too small”
- This one isn’t silent, but it is tricky: the OpenClaw versions we know reject a model whose advertised context window is too small. Increase contextWindow in the explicit declaration, and Ollama’s context accordingly.
One limitation to keep in mind once the connection is established: a properly connected local model will not necessarily behave like a large online model on long or ambiguous tasks. We publish no comparisons or throughput figures here. Start with simple, low-stakes requests, observe where the model breaks down, and keep an online provider as a fallback if you rely on the assistant every day.
#Official sources to keep handy
This guide is not based on in-house testing: it contains no duration, throughput, or score. The commands and field names follow the documentation for both projects, which changes from one version to another: ollama launch options, automatic discovery behavior, and default values. If this page differs from the documentation, the documentation takes precedence.
#Go further
The setup relies on three concepts covered in detail elsewhere on the site: the Ollama server, the context window, and tool calling.
- Install Ollama
- Installing the model server, its basic settings, and what the machine can output. https://quelllm.fr/guide/installer-ollama
- Understanding the context window
- What a token measures, why context uses memory, and how to size it. https://quelllm.fr/guide/comprendre-fenetre-contexte
- Tool calling with Ollama
- The format of tool requests and how to test them outside any agent. https://quelllm.fr/guide/appel-outil-ollama-tutoriel
- Hermes Agent with Ollama
- Another self-hosted agent connected to a local model, for comparing approaches. https://quelllm.fr/guide/hermes-agent-ollama-guide
- Install OpenClaw with Docker
- Installing and updating the gateway, and the exposure rules on a VPS. https://quelllm.fr/guide/installer-openclaw-docker
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.