Advanced 11 minAgents

OpenHands: a developer agent running on a model local

Direct response

Yes, OpenHands (formerly OpenDevin) runs on a local model through any OpenAI-compatible endpoint, but it is the most demanding agentic workload there is: below 27 to 32 billion parameters with a long context (22,000 tokens minimum, 32,768 recommended), the agent loops, misinterprets command output, and modifies the wrong file.

OpenHands gives an agent a terminal, a browser, and your repository, then lets it work: it plans, modifies files, runs commands in a sandbox, reads the output, and repeats until the task succeeds or it gives up. Running it on a local model is possible, and requires a much larger model than most people have. Here is where the line really falls, with the settings recommended by the project itself.

By Mohamed Meguedmi·Update 2026-09-28·Tested on Windows, macOS, and Linux

#What it actually does

You describe a task in French. It inspects the repository, makes a plan, then acts in a loop: run a command, read the result, modify a file, rerun the tests, read the failure, fix it. This loop is the product. It's closer to an intern with a terminal than to autocomplete.

The consequences follow. Each iteration is a complete generation over a growing context: a ten-step task costs ten long prompts. And because the agent acts instead of suggesting, its reach is everything it can access—which explains why the sandbox is part of the design, not a setting.

#The project in 2026: Agent Canvas

The Local Copilot Kit

This guide gets you to the model. The kit gets you to the coding copilot in your editor.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

OpenHands was called OpenDevin at launch, before taking its current name. The repository, maintained by All Hands AI, is released under the MIT license and has more than 89,000 stars on GitHub as of summer 2026. The project has also expanded: it no longer presents itself merely as an isolated autonomous agent, but as “the self-hosted control center for coding agents and automations,” capable of controlling OpenHands, Claude Code, Codex, or Gemini from a single interface.

In practice, the offering has been structured into several components: Agent Canvas (the control interface), a Software Agent SDK for building your own agents, an Agent Server, and an automation server for scheduled tasks. The former standalone local interface was deprecated in favor of this modular architecture; the principle remains the same for the local usage described here—an agent acting in a containerized sandbox—but component names and the launch command may differ between older documentation, including documentation you may still find indexed under the name OpenDevin.

#The model requirement, plainly stated

What you really get on an agentic development task
Model classRealistic result
7 to 8 billionFails. Poorly formed commands, misread output, loops on the same file.
14 billionSometimes succeeds at trivial tasks on a single file. Unreliable.
27 to 32 billionThe practical floor. Narrow, well-specified tasks succeed often enough to be useful.
70 billion and moreClearly better judgment, slow enough that you can start the task and go do something else.

Two capabilities matter more than benchmark scores: calling tools in the exact expected format, and stability over long contexts, since the agent rereads an ever-growing history at each step. A model that is excellent at writing a function from a prompt may be unusable in a loop. Code-oriented models generally outperform discussion models of the same size on this task.

The project's official documentation changed its recommendation during 2026: it now recommends Qwen3.6-35B-A3B as the first local model to try, an MoE (mixture-of-experts) model designed for agentic coding, with a large context and available through LM Studio, Ollama, vLLM, and SGLang. The benefit of an MoE here is direct: only 3 billion parameters are activated per token despite 35 billion total weights, making generation significantly faster than a comparably sized dense model, with VRAM requirements closer to those of a 14B model than a dense 35B.

!
Context is the hidden cost
Agent prompts include the task, plan, file contents, and command history. OpenHands documentation requires fetching a context of at least 22,000 tokens under Ollama (the default value of 4,096 is not even enough for the system prompt) and recommends 32,768. A long context costs memory before it costs time: a 32-billion-parameter model with a genuinely long context requires 24 GB or more.

#Running it locally

  1. 01
    Serve a model
    On an OpenAI-compatible endpoint, with the OLLAMA_CONTEXT_LENGTH variable raised to at least 22,000 (32,768 recommended) if you go through Ollama. This is the step people skip, and it produces agents that forget their own plan.
  2. 02
    Start the application in its containerized configuration
    It needs a container engine (Docker Desktop or Docker Engine): the agent's shell runs there, not on your host.
  3. 03
    Point it to your local access point
    With a fake API key (for example, local-llm) and a base URL reachable from the container—in Docker Desktop, http://host.docker.internal:PORT/v1 rather than localhost, which would refer to the container itself. Declare tool-calling support only if the model can actually handle it.
  4. 04
    Mount a single repository
    Ideally, a disposable clone on a branch you can delete without regret.
  5. 05
    Start with a small, verifiable task
    Fix this failing test, add this parameter, update this configuration. Then review the diff.

#The true cost of a task

An agentic task costs not one model call but a sequence of calls, each with a longer context than the last because the history accumulates. For a task that takes ten iterations, if each step adds an average of 800 tokens to the history, the tenth call is already rereading several thousand context tokens before it even produces its response. On local hardware, prompt-reading time (the “prefill”) is added to generation time on every turn, which explains why a local agent running a 32-billion-parameter model takes minutes per task, not seconds like a conventional completion.

That's the tradeoff for not paying an API bill: the cost doesn't disappear; it shifts to your graphics card and your wait time. On a poorly specified task that triggers twenty iterations because the agent is going in circles, this shifted cost quickly becomes more expensive—in electricity and time—than an equivalent API call.

#A concrete example to make things clear

Consider a realistic task: “the test test_export_csv has been failing since the last commit; fix it.” On a model in the 32-billion-parameter class, the typical workflow takes five to eight iterations: read the test and failure message, open the relevant source file, form a hypothesis about the cause, make a change, rerun the test, read the new result, and adjust if needed. Each iteration rereads the complete history of the previous ones, making the context window decisive: with only 8,000 tokens, the agent loses track after three or four steps and starts repeating hypotheses that have already been ruled out.

On a model with 7 to 8 billion parameters, the same task usually fails differently: the test is correctly identified, but the rerun command is malformed, or the model edits a neighboring file with a similar name. These are not configuration errors—they are the model's capacity limit when dealing with the strict format required by the tool-result-decision loop, regardless of the system prompt's quality. This is also why published scores on code-completion benchmarks predict very little here: a model rated highly for generating an isolated function may still be unable to chain ten coherent tool calls without going off the rails, because these two skills are not correlated in the same way depending on how the model was trained.

#Security is not optional

No secrets in the environment
The agent reads its own environment and can display it.
A disposable branch and a working clone
Never your working copy with uncommitted changes.
Restricted network access
An agent that can attach anything can exfiltrate everything it has read.
All retrieved content is hostile by default
A ticket description or README file may contain instructions intended for the agent: the mechanism is detailed in our guide to prompt injection.
Review every diff
“The tests pass” means the tests pass, not that the change is correct.

#When an assistant is better than an agent

The right tool for the task
TaskBest tool
Completion as you typeAn editor extension
A change that can be described file by fileA conversational assistant with the file open
A multi-step task with a clear success testOpenHands, if the model is large enough
Code you can't reviewNeither: you won’t be able to verify the result

In summary, local OpenHands is neither a gimmick nor a universal replacement: it is a tool to reserve for machines that can actually accommodate a model with 27 billion parameters or more and a 32,000-token context, and for tasks narrow enough that an automated test or quick review is sufficient to assess the result. Below this hardware threshold, a conventional conversational assistant with the file open remains faster and more reliable than an agent that loops without ever converging.

#FAQ

Can OpenHands run on a local model?+
Yes, through any OpenAI-compatible endpoint: Ollama, LM Studio, vLLM, or SGLang. The constraint isn’t connectivity but model capacity. This is the most demanding local workload in everyday use, and small models fail regardless of the configuration or the quality of the system prompt.
What is the minimum model?+
Realistically, a code-oriented model with 27 to 32 billion parameters, a long context, and reliable tool calling. The official documentation now recommends Qwen3.6-35B-A3B, an agentic MoE, as a first attempt. At 14 billion parameters, only trivial tasks sometimes succeed; below 8 billion, it does not work usefully.
What context should you configure under Ollama?+
At least 22,000 tokens: OpenHands documentation explicitly states that the default value of 4,096 tokens is not even enough to fit the agent's system prompt alone. It recommends 32,768 tokens via the OLLAMA_CONTEXT_LENGTH variable for comfortable use on multi-step tasks without silently truncating the history.
Is it safe to run it on my machine?+
Only with its containerized sandbox, a disposable clone of the repository, no secrets in the environment, and network access restricted to what the task actually requires. The agent runs commands it wrote itself, based on content—tickets, READMEs, dependencies—that must be treated as potentially hostile by default.
How much VRAM do you need?+
Enough to accommodate a 32-billion-parameter-class model plus a long context—24 GB or more in practice for a typical dense model. A little less is enough for an MoE such as Qwen3.6-35B-A3B, where only a fraction of the weights is activated for each generated token. In both cases, it is the context, not just the weights, that most often surprises beginners.
Why does the agent repeat the same step?+
Usually, a model that cannot produce the exact tool-call format expected by the framework, or a saturated context that silently truncated the beginning of the plan and the history of commands already executed. The fix, in order: first increase the allocated context, then switch to a larger model only if the problem persists.
Is OpenHands still the same project as OpenDevin?+
Yes, it is the same project, simply renamed early on. In 2026, it expanded well beyond the isolated autonomous agent: it now presents itself as a self-hosted control center that can also operate Claude Code, Codex, or Gemini through a modular architecture called Agent Canvas, with its own agent SDK.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.