Advanced 12 minAgents

Goose (Block): the local AI agent in your terminal

Direct response

Yes: run goose configure, choose Ollama, leave http://localhost:11434 at its default, and specify a Ollama tool-calling-compatible model. Two settings matter locally: raise OLLAMA_CONTEXT_LENGTH above the default 4096 tokens, and enable the tool shim (GOOSE_TOOLSHIM=true) if the model explains tools in text instead of calling them. Goose is now a project of the Agentic AI Foundation, not just Block.

Goose is a command-line and desktop AI agent, initially released by Block and later transferred to the Agentic AI Foundation. This guide covers configuring it with a local Ollama, the tool shim that compensates for the lack of native tool calling in some models, permission modes (autonomous by default), and the real limitations of a small local model when faced with its MCP extensions.

By Mohamed Meguedmi·Update 2026-09-28·Tested on Windows, macOS, and Linux

#What Goose is today

Goose presents itself as a native open-source AI agent—desktop application, CLI, and API—for code, workflows, and beyond, written in Rust. The repository (now aaif-goose/goose) has nearly 55,000 stars as of September 28, 2026, with version v1.52.0 released on September 23, 2026.

A fact to correct if you have read older presentations: Goose is no longer an isolated Block project; it is now part of the Agentic AI Foundation (AAIF), hosted by the Linux Foundation. Block remains the project’s origin, but its governance has changed—a detail that matters when assessing the project’s longevity before building a workflow on it.

Goose works with more than 15 model providers (Anthropic, OpenAI, Google, Ollama, OpenRouter, Azure, Bedrock, and others) and connects to more than 70 extensions through the open MCP (Model Context Protocol).

There is also a gateway to Ramalama, a local engine that serves models in OCI artifact format rather than Ollama’s proprietary format. Because its API is compatible, Goose can use it directly through its Ollama provider, with no custom code—a useful option on infrastructure already built around standard containerization tools (Podman, Docker) rather than around Ollama.

#Connect Ollama

The Local Agents Kit

Agents that act on your machine: agentic Cline, MCP, n8n + Ollama, local automations.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

Configuration goes through the interactive assistant goose configure, which asks for the provider and then the host. For Ollama, if no host is provided, Goose uses localhost:11434 by default; the http:// prefix is added automatically if the scheme is not specified.

Terminal
ollama run qwen2.5
# dans un second terminal
goose configure

For a Ollama running on another machine on the network, you must explicitly set OLLAMA_HOST=http://{hôte}:{port} before launching the configuration. For models hosted on ollama.com rather than locally, choose Ollama Cloud, not Ollama.

i
Recommended model for beginners
The official documentation uses qwen2.5 as an example of a model to launch before configuring Goose, emphasizing that it must be advertised as tool-calling compatible—not just any general-purpose chat model.

#The 4,096-token context trap

The default Ollama context window is 4096 tokens, and it silently truncates rather than returning an explicit error. With an agent such as Goose that loads project instructions (.goosehints), conversation history, and extension definitions, this limit is reached quickly.

!
Typical symptom
If Goose ignores your .goosehints files or seems to lose track with active extensions, the official documentation points first to a default context that is too short; increase it through the OLLAMA_CONTEXT_LENGTH environment variable before looking elsewhere.
Extended context
export OLLAMA_CONTEXT_LENGTH=32768

#Tool shim: fixing tool calls

Some models lack native tool-calling support, or switch mid-session to plain-text output instead of a structured call. Goose’s tool shim detects these text formats and converts them into executable tool calls. The project marks this feature as experimental.

Enable the tool shim
export GOOSE_TOOLSHIM=true
ollama pull mistral-nemo

The tool shim relies on an interpreter model separate from the main conversation model—mistral-nemo by default via Ollama, replaceable with GOOSE_TOOLSHIM_OLLAMA_MODEL. The documentation explicitly cites local models (Ollama, llama.cpp) without native tool calling as the primary use case, as well as models that mix reasoning tags (“think”) with tool calls, a common source of parsing failures.

An alternative mode uses Goose’s built-in local inference backend instead of a separate Ollama instance, through GOOSE_TOOLSHIM_BACKEND=local and a mandatory model name (GOOSE_TOOLSHIM_MODEL)—otherwise startup fails.

#Permission modes: autonomous by default

Goose offers four permission modes: fully autonomous (modifies and deletes files without confirmation), manual approval (requests confirmation for every tool), smart approval (automatically approves low-risk actions), and conversation-only mode (no modifications, no tools).

!
Standalone mode is active immediately after installation
The official documentation is explicit: autonomous mode (Autonomous Mode) is enabled by default. On a system where Goose has access to extensions capable of deleting files or running commands, explicitly changing modes before the first session is a precaution you should not skip.

You can switch modes at any time, including during a session, using /mode auto, /mode smart_approve, /mode approve, or /mode chat in the CLI, or from the bottom menu in the desktop app.

#MCP extensions and allowlist

Goose connects to extensions through the MCP protocol and installs any requested MCP server by default. For professional use, the project provides an allowlist: a YAML file hosted at a URL and referenced by the GOOSE_ALLOWLIST variable, limiting installable extensions to an explicit list of identifiers and commands.

Without this allowlist, nothing technically prevents the agent from installing an unverified third-party MCP server if the user (or the model, in autonomous mode) requests it — a point to consider together with the permission mode chosen above.

The allowlist is deployed as a simple YAML file listing authorized identifier/command pairs, hosted at a URL that Goose rereads on every restart via the GOOSE_ALLOWLIST variable. This is designed for enterprise deployments, where an administrator wants to restrict installable extensions to a preapproved list rather than trust each user or the model operating autonomously.

  1. 01
    Install the Goose CLI
    curl -fsSL https://github.com/aaif-goose/goose/releases/download/stable/download_cli.sh | bash, ou télécharger l'application de bureau depuis la documentation officielle.
  2. 02
    Prepare the Ollama model
    Verify that Ollama is running on port 11434 and launch a model explicitly advertised as tool-calling compatible before configuring Goose.
  3. 03
    Run goose configure
    Choose Ollama as the provider, confirm the proposed default host (localhost:11434), and enter the exact name of the loaded model.
  4. 04
    Increase the context if necessary
    If the agent ignores extensions or a .goosehints file, set OLLAMA_CONTEXT_LENGTH to a value higher than 4096 before restarting a session.
  5. 05
    Check the permission mode
    Before the first session with sensitive extensions, explicitly switch to manual or intelligent approval mode if the default autonomous mode is not desired.

#Which model for the available memory

The official documentation is unequivocal about a point that is often downplayed: Goose relies heavily on tool calling, and a model that doesn't support it can only handle simple conversation—all Goose extensions must then be disabled. Model choice is therefore not just a matter of response quality: it's a binary requirement for MCP extensions to work at all.

Memory scale (Q4, weights only) and realistic usage with Goose
Model sizeApproximate VRAMRealistic use with extensions
7-8B≈ 5 GBUnstable tool calling on tasks chaining multiple tools; test with GOOSE_TOOLSHIM before concluding that the model has failed
14B≈ 9 GBCommon use case for a development workstation; check the tools tag on ollama.com before loading the model
32B≈ 19–20 GBMore reliable on long tool sequences, at the cost of longer generation times on consumer GPUs
70B≈ 40 GBReserved for systems with lots of VRAM or unified memory (Mac Studio, multi-GPU workstations); rarely relevant for everyday local use

These rough figures do not replace checking the tools tag on the model page: an unannounced tool-calling-compatible variant fails on extensions regardless of its size, exactly as the official documentation cited above explains.

#Security checklist before a shared deployment

Three settings, already documented separately in this guide, together form the foundation of a safer Goose deployment than a default installation on a shared workstation or team server.

Checklist before opening Goose to other users or sensitive extensions
CheckpointSetting to verify
Permission modeSwitch from fully autonomous to manual or intelligent approval if deleting files without confirmation is not desired
Installable extensionsSet GOOSE_ALLOWLIST to a hosted YAML file that restricts the permitted identifiers and commands
The model's actual capacityConfirm the tools tag on ollama.com before connecting extensions; an incompatible model requires you to disable all of them
Enough contextRaise OLLAMA_CONTEXT_LENGTH above 4096 to prevent the security instructions themselves (.goosehints) from being silently truncated
!
The allowlist does not replace choosing the mode
A properly configured allowlist limits installable extensions, but it does not prevent an autonomous agent from using already-authorized extensions without confirmation. The two settings complement each other; they are not substitutes.

#Limits of a local model for Goose

What breaks first with a small local model
SymptomDocumented probable cause
Extensions ignored, .goosehints not trackedDefault context of 4096 tokens is too short: increase OLLAMA_CONTEXT_LENGTH
Tool calls that stop partway through a sessionThe model switches to text output: enable GOOSE_TOOLSHIM
Slow tool-shim interpreterInterpreter model too heavy: switch to a smaller model via GOOSE_TOOLSHIM_OLLAMA_MODEL
Reasoning mixed with tool callsStray « think » tags: the tool shim filters them automatically once enabled

The native DeepSeek-R1 does not support tool calling according to the official documentation, which instead suggests a community version adapted for Goose— a concrete example of the gap between a model reputed to be powerful in conversation and its actual ability to operate tools.

This gap between reasoning power and execution reliability is the structural limitation to keep in mind before installing Goose locally: a model that answers open-ended questions well offers no guarantee that it can chain tool calls without format errors. Testing a simple, verifiable task first—reading a file or running a harmless command—remains the fastest way to identify a poorly suited model before assigning the agent a multistep task.

Frequently asked questions
Is Goose still developed by Block?+
Block originated the project, but Goose is now part of the Agentic AI Foundation, hosted by the Linux Foundation. The repository has also moved between GitHub organizations, from block/goose to aaif-goose/goose—a governance change to know before building a team workflow on it, even though Block remains involved in the project.
Why does Goose ignore my extensions or my .goosehints file with Ollama?+
The most common cause is Ollama's default context, limited to 4096 tokens and silently truncated instead of returning an explicit error. The official documentation recommends increasing it through the OLLAMA_CONTEXT_LENGTH environment variable, for example to 32768, before investigating another cause.
What should I do if my local model doesn't call tools in Goose?+
Enable the tool shim with GOOSE_TOOLSHIM=true. This experimental feature detects tool calls written as plain text by models without native support and converts them into executable calls through a separate interpreter model (mistral-nemo by default, replaceable via GOOSE_TOOLSHIM_OLLAMA_MODEL). If the model mixes reasoning tags into tool calls, the tool shim filters those out automatically once enabled.
Can Goose delete files without asking for confirmation?+
Yes, in default mode: the fully autonomous mode, active upon installation according to the official documentation, allows Goose to modify and delete files without approval. Switching to manual or smart approval mode changes this behavior; do so before the first session with sensitive extensions.
Do you need a powerful GPU to run Goose locally?+
Goose itself is a lightweight Rust client; most of the load depends on the model selected through Ollama or another local provider. The tool-shim interpretation model adds a separate inference load on top of the conversation model, so choose a smaller one if responses become slow.
Can a local model without tool calling still work with Goose?+
Yes, but only for chatting: the official documentation states that Goose relies heavily on tool calling, and a model that doesn't support it can only handle simple conversation, with all extensions having to be disabled. Checking the tools tag on ollama.com before loading a model avoids this limitation.
How can you restrict the MCP extensions Goose can install?+
Set the GOOSE_ALLOWLIST variable to a hosted YAML file listing the identifiers and commands of authorized extensions; Goose rereads it on every restart. This is designed for enterprise deployments and should be combined with a non-autonomous permission mode rather than used on its own.

Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.