Agent Zero With a Local Ollama Endpoint
Agent Zero (agent0ai) gives its agent a full Linux desktop inside a Docker container. Here is the exact address for reaching a host-side Ollama, which model role needs real capability, and why self-hosting the container doesn't by itself make the browser or the model private.
Key takeaways
- Agent Zero's shipped default for host-side Ollama is
http://host.docker.internal:11434— notlocalhost, which inside the container refers to the container itself. - The agent's computer is not the inference machine. Agent Zero's Docker container (desktop, browser, orchestration) is separate from wherever Ollama actually runs the model.
- Agent Zero splits model roles: chat, utility (memory/summarization), and embedding. The project's own docs warn that very small utility models (around 4B) often fail silently at memory consolidation.
- Small local models frequently explain a command instead of calling the tool — Agent Zero ships a documented, prompt-only fix (the Tiny Local profile) rather than a code change.
- Self-hosting the container does not guarantee the built-in browser or the model provider are offline — that depends entirely on which provider and network settings you choose.
What Agent Zero actually is
Agents that act on your machine: agentic Cline, MCP, n8n + Ollama, local automations.
- Lifetime online access
- PDF + files
- 30-day refund
Agent Zero (repository agent0ai/agent-zero) bills itself as a framework that gives the agent a full Linux computer: an XFCE desktop inside a Docker container, a browser with DOM annotation, live cowork on Markdown files, spreadsheets and presentations, and a hub of 100+ community plugins. The repository had more than 19,000 GitHub stars as of September 28, 2026, with the latest tagged release, v2.13, published September 23, 2026.
It is not a code-specialist agent in the OpenCode or Cline sense. Agent Zero targets cross-cutting tasks — research, driving desktop apps, reviewing web interfaces, analysis — at the cost of a heavier footprint than a single CLI binary.
The project also ships agent profiles (on-the-fly specialists, e.g. a cautious financial-analysis persona with spreadsheet deliverables), per-project isolation of files, secrets, memories and repositories, and multi-agent cooperation where a lead agent delegates research, coding, analysis or review to focused subagents. Each of these features draws on the model's context and instruction-following capacity, which directly affects model choice once you go beyond plain chat.
Installing the Docker container
With Docker already installed, the fastest path is a single command that mounts the instance's data into a persistent volume. The project also offers a scripted installer (A0 Install) for servers and headless setups, and a graphical launcher (A0 Launcher) that manages Docker, instances and ports for you.
docker run -p 80:80 -v a0_usr:/a0/usr agent0ai/agent-zero
Once the container is running, the web UI opens on the chosen port (80 by default, or whatever you passed) to configure the model provider before the first task. On a headless server, the script installer accepts a quiet mode that creates one instance on a given port without opening any menu — useful for scripted deployment or an SSH-only remote box.
Connecting a local Ollama endpoint
The official installation guide is specific: in the chat model, utility model or embedding model settings, pick Ollama as the provider, enter the model name the way Ollama expects it (e.g. llama3.2 or qwen2.5:7b), then set the API base URL only if it differs from the default.
The default address is not localhost. Agent Zero ships Docker-friendly defaults for host-side Ollama at http://host.docker.internal:11434. This trips up users who try http://localhost:11434 from inside the container and get nothing — inside the container, localhost refers to the container itself, not the host machine.
If the container can't reach Ollama, the docs note that port 11434 must be reachable from the container; the shipped Docker Compose file maps host.docker.internal to the host gateway on Linux Docker, and if both services share the same Docker network, the address can instead be the container's name on port 11434.
Which model, for which role
Agent Zero splits models by role: chat (main conversation), utility (memory organization, summarization), and embedding (memory search). Naming format depends on the provider — with Ollama it's the bare model name (gpt-oss:20b), unlike OpenRouter which mostly requires a provider prefix.
| Role | Documented requirement |
|---|---|
| Chat (conversation) | Must follow the agent's communication format (JSON with tool_name/tool_args) |
| Utility (memory, summaries) | Very small models (e.g. 4B) often fail here; 70B-class or strong cloud flash/mini models work best |
| Embedding | Less size-sensitive; only indexes memory for similarity search |
The most commonly overlooked point: a perfectly capable chat model can still fail on the utility role if it's too small, quietly degrading memory quality without any error message — the agent just seems to "forget" or mix up facts across sessions.
Tool calling with a small local model
Agent Zero's own docs document a common failure with lightweight local models: the model explains the command instead of calling the tool that would run it. The project's fix is prompt-only, no code change: the Tiny Local profile, which keeps the standard tool-call shape but strips visible reasoning fields from the communication prompt, asking only for one executable JSON object with tool_name and tool_args.
Warning sign to watch for. If the agent narrates what it "will do" instead of acting, or repeats a malformed-message warning without correcting itself, that's the documented signature of a local model too weak for the default communication format — not a framework bug.
An alternative that keeps your current profile is adding a dedicated prompt file (a .promptinclude.md) that restates the discipline: the visible reply must be exactly one JSON object with exactly the fields tool_name and tool_args, no markdown fencing, no text before or after. The documentation is explicit about the limit of this fix: it's a prompt-and-profile adjustment, not an execution-code change — if a specific model still fails despite the discipline, the next step is switching models, not piling on more prompt text.
What the Linux desktop actually changes
Agent Zero's distinguishing feature next to terminal-only agents is that full XFCE desktop inside the container: the agent can open real GUI software (Blender to model an object, a file manager, a visible terminal window), and you watch every action, sharing the same virtual mouse and keyboard with the ability to intervene at any moment.
The built-in browser goes beyond simple page control: Annotate mode turns any webpage into a clickable instruction surface — change an element, inspect it, lift it for reuse in your own project, or leave a targeted comment the agent then treats as a task.
Self-hosting proves nothing about browser privacy. The container running on your own machine does not guarantee the built-in browser or the model's inference stay offline — that depends on the model provider you chose (local Ollama vs. a cloud API) and the container's network configuration.
Limits worth knowing before you commit
- Misleading search volume. "Agent Zero" also names unrelated entertainment content in search engines; real demand for the development framework is narrower than raw keyword volume suggests.
- Disk and RAM footprint. The container ships a full desktop on top of the agent runtime — budget more resources than a minimal CLI agent, even though inference itself runs elsewhere via Ollama.
- Undersized utility model. A too-small utility model degrades memory with no explicit error — check this first if the agent seems to lose the thread between sessions.
- Container-internal networking.
host.docker.internalworks for Docker Desktop and the provided Linux setup, but a custom deployment (custom Docker network, remote hosting) needs the API URL re-entered manually.
Related reading and sources
For the base Ollama-in-Docker setup this guide builds on, see Ollama in Docker: The Right Way. If tool-calling reliability is the main concern regardless of the agent, Ollama + MCP: Give Your Local Model Tools and CrewAI With a Local LLM cover comparable ground with different frameworks.
Primary sources used in this guide: the Agent Zero GitHub repository, its installation guide, and its local-model tool-use guide.
Frequently asked questions
What Ollama address should Agent Zero use in Docker?
http://host.docker.internal:11434 is the project's shipped default when Ollama runs on the host machine and Agent Zero runs in its container. localhost:11434 does not work from inside the container, since it points at the container itself.
Does Agent Zero need a dedicated GPU?
The container itself (desktop, browser, orchestration) needs no GPU. A GPU only matters for model inference, handled separately by Ollama or another provider, local or cloud.
Why won't my local model call tools in Agent Zero?
This is a documented behavior with small local models, which explain the command instead of executing it. The project recommends the Tiny Local profile, which simplifies the expected call format for these models.
Does the Linux desktop make Agent Zero safer than a terminal-only agent?
It contains actions inside a Docker container, limiting their reach on the host machine, but it does not guarantee inference privacy or browser privacy — both depend on the model provider and network configuration you choose.
Are Agent Zero and Hermes Agent the same thing?
No. Agent Zero targets a full computer with desktop and browser for cross-cutting tasks; Hermes Agent is a terminal-and-messaging agent centered on persistent memory and scheduled automations. Both accept a local Ollama endpoint.
A current option for local AI: GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395). Match memory to your model and software. A mini PC is a complete PC alternative; Mac/MLX and CUDA instructions require compatible hardware.
Amazon Check GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395) price →As an Amazon Associate, BestLLMfor earns from qualifying purchases, at no extra cost to you. It does not influence our independent rankings.
Found an error or have feedback? Let us know — it helps everyone who reads this guide.