Goose (Block): the local AI agent in your terminal
Yes: run goose configure, choose Ollama, leave http://localhost:11434 at its default, and specify a Ollama tool-calling-compatible model. Two settings matter locally: raise OLLAMA_CONTEXT_LENGTH above the default 4096 tokens, and enable the tool shim (GOOSE_TOOLSHIM=true) if the model explains tools in text instead of calling them. Goose is now a project of the Agentic AI Foundation, not just Block.
Goose is a command-line and desktop AI agent, initially released by Block and later transferred to the Agentic AI Foundation. This guide covers configuring it with a local Ollama, the tool shim that compensates for the lack of native tool calling in some models, permission modes (autonomous by default), and the real limitations of a small local model when faced with its MCP extensions.
#What Goose is today
Goose presents itself as a native open-source AI agent—desktop application, CLI, and API—for code, workflows, and beyond, written in Rust. The repository (now aaif-goose/goose) has nearly 55,000 stars as of September 28, 2026, with version v1.52.0 released on September 23, 2026.
A fact to correct if you have read older presentations: Goose is no longer an isolated Block project; it is now part of the Agentic AI Foundation (AAIF), hosted by the Linux Foundation. Block remains the project’s origin, but its governance has changed—a detail that matters when assessing the project’s longevity before building a workflow on it.
Goose works with more than 15 model providers (Anthropic, OpenAI, Google, Ollama, OpenRouter, Azure, Bedrock, and others) and connects to more than 70 extensions through the open MCP (Model Context Protocol).
There is also a gateway to Ramalama, a local engine that serves models in OCI artifact format rather than Ollama’s proprietary format. Because its API is compatible, Goose can use it directly through its Ollama provider, with no custom code—a useful option on infrastructure already built around standard containerization tools (Podman, Docker) rather than around Ollama.
#Connect Ollama
Agents that act on your machine: agentic Cline, MCP, n8n + Ollama, local automations.
- Lifetime online access
- PDF + files
- Lifetime updates
Configuration goes through the interactive assistant goose configure, which asks for the provider and then the host. For Ollama, if no host is provided, Goose uses localhost:11434 by default; the http:// prefix is added automatically if the scheme is not specified.
For a Ollama running on another machine on the network, you must explicitly set OLLAMA_HOST=http://{hôte}:{port} before launching the configuration. For models hosted on ollama.com rather than locally, choose Ollama Cloud, not Ollama.
#The 4,096-token context trap
The default Ollama context window is 4096 tokens, and it silently truncates rather than returning an explicit error. With an agent such as Goose that loads project instructions (.goosehints), conversation history, and extension definitions, this limit is reached quickly.
#Tool shim: fixing tool calls
Some models lack native tool-calling support, or switch mid-session to plain-text output instead of a structured call. Goose’s tool shim detects these text formats and converts them into executable tool calls. The project marks this feature as experimental.
The tool shim relies on an interpreter model separate from the main conversation model—mistral-nemo by default via Ollama, replaceable with GOOSE_TOOLSHIM_OLLAMA_MODEL. The documentation explicitly cites local models (Ollama, llama.cpp) without native tool calling as the primary use case, as well as models that mix reasoning tags (“think”) with tool calls, a common source of parsing failures.
An alternative mode uses Goose’s built-in local inference backend instead of a separate Ollama instance, through GOOSE_TOOLSHIM_BACKEND=local and a mandatory model name (GOOSE_TOOLSHIM_MODEL)—otherwise startup fails.
#Permission modes: autonomous by default
Goose offers four permission modes: fully autonomous (modifies and deletes files without confirmation), manual approval (requests confirmation for every tool), smart approval (automatically approves low-risk actions), and conversation-only mode (no modifications, no tools).
You can switch modes at any time, including during a session, using /mode auto, /mode smart_approve, /mode approve, or /mode chat in the CLI, or from the bottom menu in the desktop app.
#MCP extensions and allowlist
Goose connects to extensions through the MCP protocol and installs any requested MCP server by default. For professional use, the project provides an allowlist: a YAML file hosted at a URL and referenced by the GOOSE_ALLOWLIST variable, limiting installable extensions to an explicit list of identifiers and commands.
Without this allowlist, nothing technically prevents the agent from installing an unverified third-party MCP server if the user (or the model, in autonomous mode) requests it — a point to consider together with the permission mode chosen above.
The allowlist is deployed as a simple YAML file listing authorized identifier/command pairs, hosted at a URL that Goose rereads on every restart via the GOOSE_ALLOWLIST variable. This is designed for enterprise deployments, where an administrator wants to restrict installable extensions to a preapproved list rather than trust each user or the model operating autonomously.
- 01Install the Goose CLIcurl -fsSL https://github.com/aaif-goose/goose/releases/download/stable/download_cli.sh | bash, ou télécharger l'application de bureau depuis la documentation officielle.
- 02Prepare the Ollama modelVerify that Ollama is running on port 11434 and launch a model explicitly advertised as tool-calling compatible before configuring Goose.
- 03Run goose configureChoose Ollama as the provider, confirm the proposed default host (localhost:11434), and enter the exact name of the loaded model.
- 04Increase the context if necessaryIf the agent ignores extensions or a .goosehints file, set OLLAMA_CONTEXT_LENGTH to a value higher than 4096 before restarting a session.
- 05Check the permission modeBefore the first session with sensitive extensions, explicitly switch to manual or intelligent approval mode if the default autonomous mode is not desired.
#Which model for the available memory
The official documentation is unequivocal about a point that is often downplayed: Goose relies heavily on tool calling, and a model that doesn't support it can only handle simple conversation—all Goose extensions must then be disabled. Model choice is therefore not just a matter of response quality: it's a binary requirement for MCP extensions to work at all.
| Model size | Approximate VRAM | Realistic use with extensions |
|---|---|---|
| 7-8B | ≈ 5 GB | Unstable tool calling on tasks chaining multiple tools; test with GOOSE_TOOLSHIM before concluding that the model has failed |
| 14B | ≈ 9 GB | Common use case for a development workstation; check the tools tag on ollama.com before loading the model |
| 32B | ≈ 19–20 GB | More reliable on long tool sequences, at the cost of longer generation times on consumer GPUs |
| 70B | ≈ 40 GB | Reserved for systems with lots of VRAM or unified memory (Mac Studio, multi-GPU workstations); rarely relevant for everyday local use |
These rough figures do not replace checking the tools tag on the model page: an unannounced tool-calling-compatible variant fails on extensions regardless of its size, exactly as the official documentation cited above explains.
#Security checklist before a shared deployment
Three settings, already documented separately in this guide, together form the foundation of a safer Goose deployment than a default installation on a shared workstation or team server.
| Checkpoint | Setting to verify |
|---|---|
| Permission mode | Switch from fully autonomous to manual or intelligent approval if deleting files without confirmation is not desired |
| Installable extensions | Set GOOSE_ALLOWLIST to a hosted YAML file that restricts the permitted identifiers and commands |
| The model's actual capacity | Confirm the tools tag on ollama.com before connecting extensions; an incompatible model requires you to disable all of them |
| Enough context | Raise OLLAMA_CONTEXT_LENGTH above 4096 to prevent the security instructions themselves (.goosehints) from being silently truncated |
#Limits of a local model for Goose
| Symptom | Documented probable cause |
|---|---|
| Extensions ignored, .goosehints not tracked | Default context of 4096 tokens is too short: increase OLLAMA_CONTEXT_LENGTH |
| Tool calls that stop partway through a session | The model switches to text output: enable GOOSE_TOOLSHIM |
| Slow tool-shim interpreter | Interpreter model too heavy: switch to a smaller model via GOOSE_TOOLSHIM_OLLAMA_MODEL |
| Reasoning mixed with tool calls | Stray « think » tags: the tool shim filters them automatically once enabled |
The native DeepSeek-R1 does not support tool calling according to the official documentation, which instead suggests a community version adapted for Goose— a concrete example of the gap between a model reputed to be powerful in conversation and its actual ability to operate tools.
This gap between reasoning power and execution reliability is the structural limitation to keep in mind before installing Goose locally: a model that answers open-ended questions well offers no guarantee that it can chain tool calls without format errors. Testing a simple, verifiable task first—reading a file or running a harmless command—remains the fastest way to identify a poorly suited model before assigning the agent a multistep task.
- OpenCode + Ollama: a coding agent in your terminal
- Cline + Ollama: 100% local coding agent in VS Code
- MCP: what is it? The Model Context Protocol explained
- Tool calling with Ollama: tutorial
- Source: the official README for the Goose repository
- Source: official provider documentation
- Source: official tool shim documentation
Is Goose still developed by Block?+
Why does Goose ignore my extensions or my .goosehints file with Ollama?+
What should I do if my local model doesn't call tools in Goose?+
Can Goose delete files without asking for confirmation?+
Do you need a powerful GPU to run Goose locally?+
Can a local model without tool calling still work with Goose?+
How can you restrict the MCP extensions Goose can install?+
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.