OpenHands: a developer agent running on a model local
Yes, OpenHands (formerly OpenDevin) runs on a local model through any OpenAI-compatible endpoint, but it is the most demanding agentic workload there is: below 27 to 32 billion parameters with a long context (22,000 tokens minimum, 32,768 recommended), the agent loops, misinterprets command output, and modifies the wrong file.
OpenHands gives an agent a terminal, a browser, and your repository, then lets it work: it plans, modifies files, runs commands in a sandbox, reads the output, and repeats until the task succeeds or it gives up. Running it on a local model is possible, and requires a much larger model than most people have. Here is where the line really falls, with the settings recommended by the project itself.
#What it actually does
You describe a task in French. It inspects the repository, makes a plan, then acts in a loop: run a command, read the result, modify a file, rerun the tests, read the failure, fix it. This loop is the product. It's closer to an intern with a terminal than to autocomplete.
The consequences follow. Each iteration is a complete generation over a growing context: a ten-step task costs ten long prompts. And because the agent acts instead of suggesting, its reach is everything it can access—which explains why the sandbox is part of the design, not a setting.
#The project in 2026: Agent Canvas
This guide gets you to the model. The kit gets you to the coding copilot in your editor.
- Lifetime online access
- PDF + files
- Lifetime updates
OpenHands was called OpenDevin at launch, before taking its current name. The repository, maintained by All Hands AI, is released under the MIT license and has more than 89,000 stars on GitHub as of summer 2026. The project has also expanded: it no longer presents itself merely as an isolated autonomous agent, but as “the self-hosted control center for coding agents and automations,” capable of controlling OpenHands, Claude Code, Codex, or Gemini from a single interface.
In practice, the offering has been structured into several components: Agent Canvas (the control interface), a Software Agent SDK for building your own agents, an Agent Server, and an automation server for scheduled tasks. The former standalone local interface was deprecated in favor of this modular architecture; the principle remains the same for the local usage described here—an agent acting in a containerized sandbox—but component names and the launch command may differ between older documentation, including documentation you may still find indexed under the name OpenDevin.
#The model requirement, plainly stated
| Model class | Realistic result |
|---|---|
| 7 to 8 billion | Fails. Poorly formed commands, misread output, loops on the same file. |
| 14 billion | Sometimes succeeds at trivial tasks on a single file. Unreliable. |
| 27 to 32 billion | The practical floor. Narrow, well-specified tasks succeed often enough to be useful. |
| 70 billion and more | Clearly better judgment, slow enough that you can start the task and go do something else. |
Two capabilities matter more than benchmark scores: calling tools in the exact expected format, and stability over long contexts, since the agent rereads an ever-growing history at each step. A model that is excellent at writing a function from a prompt may be unusable in a loop. Code-oriented models generally outperform discussion models of the same size on this task.
The project's official documentation changed its recommendation during 2026: it now recommends Qwen3.6-35B-A3B as the first local model to try, an MoE (mixture-of-experts) model designed for agentic coding, with a large context and available through LM Studio, Ollama, vLLM, and SGLang. The benefit of an MoE here is direct: only 3 billion parameters are activated per token despite 35 billion total weights, making generation significantly faster than a comparably sized dense model, with VRAM requirements closer to those of a 14B model than a dense 35B.
#Running it locally
- 01Serve a modelOn an OpenAI-compatible endpoint, with the OLLAMA_CONTEXT_LENGTH variable raised to at least 22,000 (32,768 recommended) if you go through Ollama. This is the step people skip, and it produces agents that forget their own plan.
- 02Start the application in its containerized configurationIt needs a container engine (Docker Desktop or Docker Engine): the agent's shell runs there, not on your host.
- 03Point it to your local access pointWith a fake API key (for example, local-llm) and a base URL reachable from the container—in Docker Desktop, http://host.docker.internal:PORT/v1 rather than localhost, which would refer to the container itself. Declare tool-calling support only if the model can actually handle it.
- 04Mount a single repositoryIdeally, a disposable clone on a branch you can delete without regret.
- 05Start with a small, verifiable taskFix this failing test, add this parameter, update this configuration. Then review the diff.
- Serving a model with a long context
- Have a local model review code
- Cline: the coding agent built into the editor
- The QuelLLM kit for building a local agent
- Official documentation: local models with OpenHands
- Official OpenHands repository on GitHub
- Independent review of OpenHands (2026)
#The true cost of a task
An agentic task costs not one model call but a sequence of calls, each with a longer context than the last because the history accumulates. For a task that takes ten iterations, if each step adds an average of 800 tokens to the history, the tenth call is already rereading several thousand context tokens before it even produces its response. On local hardware, prompt-reading time (the “prefill”) is added to generation time on every turn, which explains why a local agent running a 32-billion-parameter model takes minutes per task, not seconds like a conventional completion.
That's the tradeoff for not paying an API bill: the cost doesn't disappear; it shifts to your graphics card and your wait time. On a poorly specified task that triggers twenty iterations because the agent is going in circles, this shifted cost quickly becomes more expensive—in electricity and time—than an equivalent API call.
#A concrete example to make things clear
Consider a realistic task: “the test test_export_csv has been failing since the last commit; fix it.” On a model in the 32-billion-parameter class, the typical workflow takes five to eight iterations: read the test and failure message, open the relevant source file, form a hypothesis about the cause, make a change, rerun the test, read the new result, and adjust if needed. Each iteration rereads the complete history of the previous ones, making the context window decisive: with only 8,000 tokens, the agent loses track after three or four steps and starts repeating hypotheses that have already been ruled out.
On a model with 7 to 8 billion parameters, the same task usually fails differently: the test is correctly identified, but the rerun command is malformed, or the model edits a neighboring file with a similar name. These are not configuration errors—they are the model's capacity limit when dealing with the strict format required by the tool-result-decision loop, regardless of the system prompt's quality. This is also why published scores on code-completion benchmarks predict very little here: a model rated highly for generating an isolated function may still be unable to chain ten coherent tool calls without going off the rails, because these two skills are not correlated in the same way depending on how the model was trained.
#Security is not optional
- No secrets in the environment
- The agent reads its own environment and can display it.
- A disposable branch and a working clone
- Never your working copy with uncommitted changes.
- Restricted network access
- An agent that can attach anything can exfiltrate everything it has read.
- All retrieved content is hostile by default
- A ticket description or README file may contain instructions intended for the agent: the mechanism is detailed in our guide to prompt injection.
- Review every diff
- “The tests pass” means the tests pass, not that the change is correct.
#When an assistant is better than an agent
| Task | Best tool |
|---|---|
| Completion as you type | An editor extension |
| A change that can be described file by file | A conversational assistant with the file open |
| A multi-step task with a clear success test | OpenHands, if the model is large enough |
| Code you can't review | Neither: you won’t be able to verify the result |
In summary, local OpenHands is neither a gimmick nor a universal replacement: it is a tool to reserve for machines that can actually accommodate a model with 27 billion parameters or more and a 32,000-token context, and for tasks narrow enough that an automated test or quick review is sufficient to assess the result. Below this hardware threshold, a conventional conversational assistant with the file open remains faster and more reliable than an agent that loops without ever converging.
#FAQ
Can OpenHands run on a local model?+
What is the minimum model?+
What context should you configure under Ollama?+
Is it safe to run it on my machine?+
How much VRAM do you need?+
Why does the agent repeat the same step?+
Is OpenHands still the same project as OpenDevin?+
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.