Intermediate 11 minAgents

OpenCode + Ollama: a coding agent in your terminal

Direct response

OpenCode and Ollama are not competitors: OpenCode is the coding agent in the terminal, while Ollama is the engine that serves the model locally. To connect them, install OpenCode, declare a provider pointing to http://localhost:11434/v1 in opencode.json, or run ollama launch opencode. Set the context to at least 64,000 tokens, as required by the Ollama documentation.

This guide installs OpenCode on macOS, Linux, or Windows, connects it to Ollama using both the official and manual methods, chooses a coding model that fits on a real machine, explains context settings (the number-one pitfall) and permissions, which many people assume are stricter than they are. It also compares OpenCode with Cline and Aider.

By Mohamed Meguedmi·Update 2026-09-29·Tested on Windows, macOS, and Linux

#OpenCode with Ollama: who does what

OpenCode is an open-source coding agent that runs in the terminal: it reads your project, modifies files, and runs commands. It doesn't provide a model itself. Ollama, on the other hand, downloads and runs models on your machine and exposes a local API. The two complement each other: OpenCode sends its requests to Ollama, which responds with a local model. So the question isn't “OpenCode or Ollama”; you use both.

The value of running locally is not ideological. A coding agent sees everything: the directory tree, configuration files, and business logic. With a local model, this context stays on the machine, with no per-token billing or network dependency. The price is capacity: a local model with a few tens of billions of parameters does not match the largest hosted models, and long tasks run more slowly.

i
OpenCode also offers paid services, with no obligation
OpenCode’s documentation presents OpenCode Zen (a list of models tested by the team) and OpenCode Go (a subscription to hosted open models) as optional providers. They have nothing to do with 100% local use through Ollama.

#Prerequisites: machine, terminal, and Ollama

The Local Copilot Kit

This guide gets you to the model. The kit gets you to the coding copilot in your editor.

  • Lifetime online access
  • PDF + files
  • Lifetime updates
Ollama in service
Verify that a model responds before configuring anything: ollama list, then ollama run with a code model. If that works, only the OpenCode configuration remains.
A modern terminal
The OpenCode documentation lists WezTerm, Alacritty, Ghostty, and Kitty. On Windows, it recommends using WSL for better performance and full compatibility.
Memory for context
An agent reads files and accumulates history. Ollama indicates that OpenCode requires a context length of at least 64,000 tokens, which increases the required memory beyond the model weights.
A git repository
Not required, but recommended: a git diff or git restore cleanly undoes a failed session.

#Install OpenCode and connect it to Ollama

OpenCode can be installed with a script or package managers. The documentation lists the official script, npm, bun, pnpm, yarn, Homebrew, Arch, Chocolatey, Scoop, Mise, and Docker. For macOS and Linux with Homebrew, it recommends the official tap over the base formula, which is updated less often.

Script installation (macOS, Linux)
curl -fsSL https://opencode.ai/install | bash
Installation via npm (all platforms)
npm install -g opencode-ai
Homebrew, official tap
brew install anomalyco/tap/opencode
Windows, Chocolatey, or Scoop
choco install opencode
scoop install opencode

#Quick method: ollama launch opencode

Ollama can launch OpenCode with a selected model. The command ollama launch opencode starts OpenCode with a configuration passed on the command line, without overwriting your ~/.config/opencode/opencode.json file; your existing OpenCode settings continue to apply. With the --config option, Ollama configures OpenCode without opening an interactive session. Models defined only in opencode.json do not appear in the ollama launch selector.

Run OpenCode with a Ollama model
ollama launch opencode

#Manual method: declare Ollama as the provider

If you prefer to stay in control, add a provider to opencode.json, either globally (~/.config/opencode/opencode.json) or at the project root. The OpenAI-compatible endpoint for Ollama is http://localhost:11434/v1. Each key under models must exactly match the model name shown by ollama list.

opencode.json
{
  "$schema": "https://opencode.ai/config.json",
  "provider": {
    "ollama": {
      "npm": "@ai-sdk/openai-compatible",
      "name": "Ollama",
      "options": { "baseURL": "http://localhost:11434/v1" },
      "models": {
        "qwen3-coder:30b": { "name": "Qwen3-Coder 30B" },
        "devstral-small-2:24b": { "name": "Devstral Small 2" }
      }
    }
  }
}

Restart OpenCode: the model selector lists the models declared under the Ollama provider.

#The number-one pitfall: context length

Many initial attempts fail because the context is too short: the agent loses the beginning of the task, repeats file reads, or produces inconsistent modifications. Ollama does not set a single context length: the documentation specifies 4k tokens below 24 GB of VRAM, 32k between 24 and 48 GB, and 256k starting at 48 GB. It states that tasks requiring a large context, such as agents and coding tools, should be configured for at least 64,000 tokens.

Default context for Ollama based on VRAM, and the action to take
Available VRAMDefault contextFor OpenCode
Less than 24 GB4,000 tokensMove to at least 64,000
24 to 48 GB32,000 tokensMove to at least 64,000
48 GB or more256,000 tokensSufficient; monitor memory

There are two ways to change the value. In the Ollama app, a settings slider sets the context. On the command line, the OLLAMA_CONTEXT_LENGTH variable applies when the server starts. A larger context uses more memory: check with ollama ps that the model fits entirely on the GPU, because a model that spills onto the CPU becomes very slow.

64,000-token context at server launch
OLLAMA_CONTEXT_LENGTH=64000 ollama serve

#Which local model for a coding agent

An agent must call tools reliably: read, write, execute. Choose a model that shows tool capability in the Ollama library. Here are candidates whose tags and sizes were recorded on ollama.com.

Tool-use models suited to code agents (Ollama library, September 2026)
ModelTagDownload sizeNote
Qwen3-Coder 30Bqwen3-coder:30b19 GBAdvertised native context of 256K; tool-use capability
Devstral Small 2devstral-small-2:24b15 GBAdvertised 384K context; tools and images
Qwen3.5qwen3.5:9b, 27b, 35bby sizeTools, vision, and reasoning by size
gpt-ossgpt-oss:20b, 120bby sizeTools and reasoning

These sizes are the file sizes; add the context cache, which grows with the requested 64,000 tokens. The 15 GB model (Devstral Small 2) is the most accessible on a 24 GB machine; the 19 GB model (Qwen3-Coder 30B) requires more headroom. On a more modest machine, a smaller model can handle simple, targeted tasks, but falls behind more quickly on refactoring that touches multiple files.

!
A “Qwen3-Coder” 7B does not exist in the Ollama library
Qwen3-Coder is offered only in 30B and 480B. For a small model with tools, choose a smaller size from a general-purpose family such as Qwen3.5, and accept shorter tasks.

#A real workflow: add a route and its test

Concrete example: add a GET /health route to a small Express API and cover it with a test. OpenCode’s documentation recommends starting with /init, which analyzes the project and creates an AGENTS.md file at the root; commit it to git to help the agent understand the structure.

  1. 01
    Open the project
    Go to the repository root, launch opencode, then run /init the first time. Check at the bottom that the selected model is actually a Ollama model.
  2. 02
    Switch to Plan mode
    The Tab key switches between Plan and Build modes. Plan mode disables modifications: the agent only suggests how it would proceed. Ask it for a plan before making any changes.
  3. 03
    Describe the task
    Provide the context as you would to a junior developer: “Add a GET /health route in src/server.js that returns status ok, then add a test in test/health.test.js.” The @ character lets you search for a file in the project.
  4. 04
    Switch to Build mode
    When you are happy with the plan, press Tab again and ask it to apply the changes.
  5. 05
    Run and iterate
    The agent can run npm test, read the output, and fix the issue. This execute, observe, fix loop is the core of agentic systems.
  6. 06
    Verify and commit
    Review git diff before committing. If the session went off track, git restore returns the files to their previous state.

With a local model, every step is slower than with a hosted model, and an ambiguous instruction sometimes requires rephrasing. The benefit is control: the code and context stay on your machine, and you choose the model, its quantization, and its context size based on your hardware.

#Permissions: what the agent can do without asking you

A common assumption is that the agent changes nothing without your approval. That is not the default behavior. OpenCode's documentation states that, without configuration, most permissions are set to allow, meaning they run without asking; only a few, such as external_directory and doom_loop, are set to ask. .env files are denied read access by default, except for .env.example.

For an agent controlling a less reliable local model, a stricter setting is prudent. The opencode.json file accepts a permission section where each action can be allow, ask, or deny, including a global * rule set to ask.

Request confirmation for every action
{
  "$schema": "https://opencode.ai/config.json",
  "permission": { "*": "ask" }
}
→
Before any session on a repository that matters
Commit or set aside your changes before launching the agent, keep the * rule set to ask until you trust the model, and review the proposed shell commands. The --auto option automatically approves anything that is not explicitly denied: reserve it for a disposable repository.

#OpenCode, Cline, or Aider: which should you choose?

Three coding agents that connect to Ollama
ToolFormStrengthChoose if
OpenCodeTerminal interface, also available as an application and IDE extensionEditor-independent, multi-provider, Plan and Build modesYou live in the terminal or work on remote servers
ClineVS Code extensionDiffs displayed in the editorYour workflow revolves around VS Code
AiderCommand lineStrong Git integration, automatic commitsYou want fine-grained control over the files added to the context

All three use Ollama's API; trying the two closest to your usual workflow takes an hour and is more useful than a comparison. For inline completion rather than an agent, take a look at Tabby.

#Troubleshooting

The model does not appear
The name in opencode.json does not match ollama list. Copy the exact name, including the tag. With ollama launch, a model defined only in opencode.json does not appear in its selector.
The agent forgets what it just read
The context is too short: increase it to 64,000 tokens (OLLAMA_CONTEXT_LENGTH) and check with ollama ps.
Connection refused
The server Ollama isn’t running or isn’t listening on the default port 11434. Run ollama list, then test the baseURL.
Tool calls that fail
The model doesn't handle tools well. Choose a model marked tools in the Ollama library.
Very slow responses
The model and its context exceed the VRAM, so part of the workload runs on the CPU. Reduce the context, model size, or quantization.
Frequently asked questions about OpenCode and Ollama
How do you use OpenCode with Ollama?+
Install OpenCode, then run ollama launch opencode, or add an openai-compatible provider to opencode.json with the baseURL http://localhost:11434/v1 and your models. Set the Ollama context to at least 64,000 tokens; otherwise, the agent quickly loses track of the task. Finally, check with ollama ps that the model fits on the GPU.
What's the difference between OpenCode and Ollama?+
OpenCode is a coding agent that runs in the terminal; it reads and modifies your project. Ollama is an engine that runs language models on your machine and exposes an API. They complement each other: OpenCode sends its requests to the model served by Ollama.
Does OpenCode work on Windows with Ollama?+
Yes. OpenCode's documentation offers Chocolatey, Scoop, npm, or Docker and recommends WSL for better performance. Ollama runs natively on Windows and serves the API at localhost:11434: on native Windows, the same baseURL works. If you use WSL, make sure OpenCode can reach the Ollama server there, and install Ollama in WSL if in doubt.
What context should you set in Ollama for OpenCode?+
The Ollama documentation states that OpenCode requires at least 64,000 context tokens. The default for Ollama ranges from 4,000 tokens with less than 24 GB of VRAM to 256,000 starting at 48 GB. Set the value through the application or the OLLAMA_CONTEXT_LENGTH variable, then check the memory in use with ollama ps.
Does OpenCode modify my files without asking?+
By default, yes for most actions: the documentation indicates that most permissions are set to allow. To be prompted before every action, add a global rule * set to ask in the permission section of opencode.json, and commit your work before each session.
Which local model should you choose for OpenCode?+
Choose a model displaying the tools capability in the Ollama library: Qwen3-Coder 30B (19 GB) or Devstral Small 2 24B (15 GB) are two starting points for a machine with 24 GB or more. On a more modest machine, test a small Qwen3.5 size on short tasks.

#Go further

Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.