Intermediate 10 minIDE

Zed + Ollama: the ultra-fast editor with an assistant local

Direct response

Yes, Zed connects natively to Ollama: install Ollama, run ollama pull d's model, verify that the server is running, then choose that model from the dropdown, which Zed populates automatically. One setting to know before using it: Zed sends a context window of only 4096 tokens to Ollama by default, well below most recent models, which is enough to explain truncated responses on a large file.

Zed is a multi-user code editor written in Rust, known for its responsiveness. This guide covers configuring it with Ollama for the local AI assistant, adjusting the context window, automatically discovering models, and the permission and sandbox mechanisms that constrain what Zed’s agent can do on your machine.

By Mohamed Meguedmi·Update 2026-09-28·Tested on Windows, macOS, and Linux

#Configure Ollama in Zed

Zed's official documentation describes a four-step process: download and install Ollama, pull a model, verify that the Ollama server is running, then select that model from Zed's drop-down menu.

Terminal
ollama pull mistral
ollama serve

On macOS, launching the Ollama.app application is enough to start the server; on Linux or from a shell, the ollama serve command does the same thing. Zed is described by its publisher as a high-performance, multi-user code editor created by the authors of Atom and Tree-sitter; the latest stable version as of September 28, 2026 is v1.21.0, released on September 23, 2026, and the repository had nearly 91,000 stars on GitHub on the same date.

This Ollama configuration supports all of Zed's AI features that rely on a language model: the Agent panel for multi-file tasks, the Inline Assistant for targeted editor changes, and terminal threads that let the agent suggest commands. External agents and terminal threads may nevertheless require their own local model configuration, separate from the one described here for Zed's native features.

#Automatic model detection

The Local Copilot Kit

This guide gets you to the model. The kit gets you to the coding copilot in your editor.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

Zed automatically detects the models that Ollama has already downloaded and offers them in the selection menu. To disable this detection and list the models yourself with their exact capabilities, set auto_discover to false in the settings.

settings.json
{
  "language_models": {
    "ollama": {
      "api_url": "http://localhost:11434",
      "auto_discover": false,
      "available_models": [
        {
          "name": "qwen2.5-coder",
          "display_name": "qwen 2.5 coder",
          "max_tokens": 32768,
          "supports_tools": true,
          "supports_thinking": true,
          "supports_images": true
        }
      ]
    }
  }
}

This manual mode has practical value: you can force supports_tools or max_tokens for a model that Zed would not classify correctly automatically, or document in the project's version-controlled configuration file which model the team is expected to use.

The supports_tools field deserves particular attention: it tells Zed whether the model can receive tool definitions in the Zed Agent (file editing, command execution) rather than being limited to simple chat. A Ollama model without declared tool-calling capability will work for the inline assistant or autocompletion, but not for multi-file tasks in the Agent panel.

#The 4096-token context window

Zed passes the context length to Ollama through the num_ctx parameter, with 4096 tokens by default according to the official documentation. That’s significantly less than the native context window of most recent models (often 32,000 tokens or more), and a medium-sized file may be enough to exceed this limit.

!
Surprise gradient
This isn't a model limitation but a default setting in Zed itself. A model advertised with a 128k-token context will still produce truncated responses on a large file if Zed continues sending num_ctx=4096 and this isn't corrected in the settings.
Increase the context
{
  "language_models": {
    "ollama": {
      "context_window": 8192
    }
  }
}

This context_window setting applies to all configured Ollama models in Zed; a different context per model is defined via max_tokens in available_models, with auto_discover disabled so the value takes effect.

#llama.cpp and LM Studio, the local alternatives

Zed supports llama.cpp just like Ollama, with automatic discovery of models served in router mode, refined by an /models/sse stream that requires a recent version of the llama.cpp server. LM Studio is also supported through its local API server, started with lms server start.

Compare the three local paths in Zed
BackendModel discoverySpecial feature
OllamaAutomatic (models already pulled)Context to set manually (4096 by default)
llama.cppAutomatic in router mode (recent build required)On-demand loading with the -hf option
LM StudioManual via the loaded-model listGraphical model management interface in addition to the API

#Agent tool permissions

Since version 0.224.0, approval for Zed agent tools is configured through agent.tool_permissions.default; before that version, a simple Boolean (agent.always_allow_tool_actions, false by default) governed everything. The new system supports regex pattern rules, with three possible outcomes: allow, deny, or always ask for confirmation.

settings.json — rules per tool
{
  "agent": {
    "tool_permissions": {
      "default": "allow",
      "tools": {
        "terminal": {
          "default": "confirm",
          "always_allow": [
            { "pattern": "^cargo\\s+(build|test|check)" }
          ],
          "always_confirm": [{ "pattern": "sudo\\s+/" }]
        }
      }
    }
  }
}

This example, taken from the official documentation, automatically authorizes precise cargo commands in the terminal tool while consistently requiring confirmation for any sudo command affecting the system root—a useful level of granularity when the connected model is a small local model that is less predictable than a reference cloud model.

#The sandbox: what it really protects

Beyond declarative permissions, Zed provides an operating-system-level sandbox for its agent's tool calls. The documentation is precise about its scope: it applies only to terminal and fetch tools, not to Zed itself, language servers, extensions, tasks, or regular terminal tabs.

Terminal tool
The sandbox restricts writes to disk and outbound network access from commands launched by the agent; Git metadata is protected.
fetch tool
The sandbox restricts the hosts that the agent can actually contact.
Linux prerequisites
An executable, non-setuid bwrap binary must be present in the PATH.
Windows requirements
WSL must be available; the documentation states that its sandbox is weaker than on Linux or macOS and may not block all escape attempts.
i
Permissions and sandboxing complement each other
Pattern-based permissions limit the agent's ability to trigger an action; the sandbox, once the action has started, limits what it can actually access on the system. The two mechanisms are not substitutes for each other.

This limited scope (terminal and fetch only) is important to understand before assigning sensitive tasks to a less predictable small local model: file changes made through edit_file or write_file go through declarative permissions and agent profiles, not the sandbox’s system isolation. Human review therefore remains useful for these actions, even when the sandbox is active for everything else.

#MCP servers in Zed

Beyond Ollama as a model provider, Zed uses the Model Context Protocol to interact with external context servers: databases, ticketing systems, and corporate documentation. The official documentation states that Zed supports MCP Tools and Prompts features, giving the agent additional capabilities beyond the built-in tools (file editing, terminal).

To ensure that a specific MCP server is actually used rather than competing with Zed’s built-in tools, the documentation illustrates a dedicated agent profile (using the container-use server as an example) that disables the built-in tools and enables only those from the selected MCP server, via enable_all_context_servers set to false in the profile configuration.

settings.json — excerpt from the official MCP profile
{
  "enable_all_context_servers": false,
  "context_servers": {
    "container-use": {
      "tools": {
        "environment_create": true,
        "environment_add_service": true,
        "environment_update": true,
        "environment_run_cmd": true
      }
    }
  }
}

This per-profile granularity aligns with the tool-permission system described above: with a local Ollama model that is less predictable than a reference cloud model, limiting the tools actually exposed to the agent — whether they come from an MCP server or Zed's native capabilities — reduces the error surface before permissions and sandboxing are even considered.

#Predictive editing (Zeta) locally too

Zed offers a feature distinct from the chat agent: predictive editing, which suggests the next change as you type and accepts it with a press of Tab. The default provider is Zeta, an open-source model developed by Zed itself—distinct from the conversational models configured for the Agent panel or inline assistant.

To keep this feature fully local, the official documentation lets you configure a dedicated Ollama provider for predictive editing, with model variants intended for this use (especially zeta2), rather than reusing the chat model already configured for the agent—the two settings are independent in Zed. A smaller model designed to predict one modification at a time rather than hold a multi-turn conversation generally responds faster than a general-purpose chat model for this specific task, which matters more here than raw reasoning capacity: perceived latency while typing is the primary criterion, well ahead of response richness.

→
Two models, two distinct configurations
The model used for predictive editing (Zeta or its Ollama equivalent) and the one used for the chat agent or online assistant are configured separately. Changing num_ctx for the agent, for example, does not affect predictive editing as long as its own provider is not adjusted.

#Ollama on a remote server

If Ollama runs on another machine or requires a key (as with Ollama Turbo, the hosted version), enter the key in the provider's interface or through the OLLAMA_API_KEY variable, and point the API URL to the remote endpoint rather than localhost.

This remote setup has a practical benefit for a team: a single Ollama server sized for a reasonably large model can serve multiple Zed workstations configured with the same API URL, so each developer does not have to run and reload the model on their own machine. The context window and declared capabilities (supports_tools, supports_thinking) then need to be configured only once, in the project’s shared configuration rather than in each person’s settings.

Frequently asked questions
Why does Zed truncate responses on a large file with Ollama?+
By default, Zed sends Ollama a context window of only 4096 tokens (num_ctx parameter), regardless of the native context window advertised by the model. This is not a model limitation but a default setting in Zed itself. The context_window setting in Zed's Ollama settings lets you increase it, for example to 8192 or more depending on available VRAM.
Does Zed automatically detect my Ollama models?+
Yes, by default, Zed detects models already pulled by Ollama and offers them in the selection menu. To list models yourself with their exact capabilities—tool, reasoning, and image support—you must disable auto_discover and declare them manually in available_models, with a separate max_tokens value for each.
Does Zed's sandbox protect all of the agent's tools?+
No. The official documentation specifies that it applies only to terminal and fetch tools. Other tools, including file editing through edit_file or write_file, remain governed by declarative permissions and agent profiles, without additional system isolation—a human review is still useful for these actions.
Can you use LM Studio or llama.cpp instead of Ollama in Zed?+
Yes, all three are supported natively for Zed's AI features. llama.cpp and Ollama offer automatic discovery of already-loaded models; LM Studio requires starting its own local API server with the command lms server start before Zed can connect to it.
Is Zed's sandbox as reliable on Windows as it is on Linux?+
No, the official documentation explicitly states this: the Windows sandbox relies on WSL and is weaker than on Linux or macOS, where it relies on bwrap; it may not prevent every escape attempt on this platform, which you should keep in mind before entrusting it with a sensitive task.
Can Zed use MCP servers with a local Ollama model?+
Yes. Zed supports the Model Context Protocol's Tools and Prompts features regardless of the configured model provider, including Ollama. A dedicated agent profile lets you disable built-in tools and expose only those from a specific MCP server, limiting the error surface with a less predictable local model.
Can Zed's predictive editing run entirely locally?+
Yes. The default provider, Zeta, is an open-source model developed by Zed, but the official documentation lets you replace it with a local Ollama model (variants such as zeta2 are intended for this use). This setting is independent of the model used for the chat agent or online assistant.

Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.