BestLLMfor Your hardware. Your LLM. Your call.
◆ The kits◆ Kits APIOpen data Find my LLM
Guide · 2026-09-24

OpenHands With a Local Model

◆ Local Copilot — Replace Copilot with a code assistant that runs on your machine · $24 · or all kits $49 →

OpenHands gives an agent a terminal, a browser and your repository, then lets it work. Running it on a local model is possible and demands a bigger model than most people have — here is where the line actually sits.

By Mohamed Meguedmi·Last updated 2026-09-24·10 min read·Tested on Windows, macOS, Linux

Key takeaways

  • OpenHands is an autonomous software-engineering agent: it plans, edits files, runs commands in a sandbox, reads the output and iterates until a task is done or it gives up.
  • It is model-agnostic and can be pointed at a local endpoint — but this is the most demanding local workload in common use.
  • Below roughly 30B parameters the experience is poor: the agent loops, misreads command output and edits the wrong file. This is a capability limit, not a configuration mistake.
  • Everything runs in a container by design, which is not optional: the agent executes commands it wrote itself.
  • The realistic local pattern is narrow, well-specified tasks on a repository you can review — not "implement this feature" on a codebase you do not know.

What it does, concretely

The Local Copilot Kit

Replace GitHub Copilot and Cursor with a code assistant that runs 100% on your machine — the reference guide, configs included.

  • Lifetime online access
  • PDF + files
  • 30-day refund

Give it a task in plain language. It inspects the repository, forms a plan, and then acts in a loop: run a command, read the result, edit a file, run the tests, read the failure, fix. That loop is the product. It is closer to an intern with a terminal than to autocomplete.

The consequences follow directly. Each iteration is a full generation over a growing context, so a task that takes ten steps costs ten long prompts. And since the agent acts rather than suggests, the blast radius is whatever it can reach — which is why the sandbox is part of the design rather than a setting.

The model requirement, stated plainly

Model classRealistic outcome on an agentic coding task
7B–8BFails. Malformed commands, misread output, loops on the same file.
14BOccasionally completes trivial single-file tasks; unreliable.
27B–32BThe practical floor. Narrow, well-specified tasks succeed often enough to be useful.
70B+Noticeably better judgement; slow enough that you will start the task and leave.

Two capabilities matter more than benchmark scores: reliable tool calling in the exact format expected, and long-context stability, because the agent re-reads a growing history every step. A model that is excellent at writing a function from a prompt can still be useless in a loop. Coding-oriented models usually outperform general chat models of the same size here — see the coding model ranking.

Context is the hidden cost. Agent prompts carry the task, the plan, the file contents and the command history. Long contexts are expensive in memory before they are expensive in time — the arithmetic is in what a long context costs in VRAM. A 32B model with a genuinely long context is a 24 GB-plus proposition.

Running it locally

  1. Serve a model at an OpenAI-compatible endpoint with a long context configured — this is the step people skip, and it produces agents that forget their own plan.
  2. Start the application in its container setup. It needs a container runtime; the agent's shell lives there, not on your host.
  3. Point it at your local endpoint with a placeholder API key, and declare function-calling support only if the model really has it.
  4. Mount one repository, ideally a scratch clone on a branch you are happy to delete.
  5. Start with a small, verifiable task: fix this failing test, add this parameter, update this configuration. Then read the diff.

Safety is not optional here

  • No credentials in the environment. The agent reads its own environment and may print it.
  • A throwaway branch and a scratch clone. Never your working tree with uncommitted changes.
  • Restrict network access unless the task needs it. An agent that can reach anything can exfiltrate anything it has read.
  • Treat fetched content as hostile. An issue description or a README can carry instructions aimed at the agent — the mechanism is in prompt injection.
  • Review every diff. "The tests pass" means the tests pass, not that the change is correct.

When an assistant beats an agent

TaskBetter tool
Inline completion while you typeAn editor extension — see local coding assistants
A change you can describe file by fileA conversational assistant with the file open
A multi-step task with a clear success testOpenHands, if your model is large enough
Anything on a codebase you cannot reviewNeither — you cannot verify the result

Verdict

OpenHands is the most honest demonstration of where local models stand on agentic work: the framework is ready, and the hardware requirement is the barrier. With a 32B-class coding model, a long context and a scratch repository, it genuinely closes small tasks unattended. With anything smaller you will spend more time watching a loop than you would have spent doing the work — and no configuration fixes that.

Frequently asked questions

Can OpenHands run on a local model?

Yes, through any OpenAI-compatible endpoint. The constraint is capability: this is the most demanding common local workload, and small models fail at it regardless of configuration.

What is the minimum model for OpenHands?

Realistically a 27B–32B coding-oriented model with a long context. 14B occasionally completes trivial tasks; 7B–8B does not work.

Is it safe to run on my own machine?

Only with its container sandbox, a scratch clone, no credentials in the environment and restricted network access. The agent executes commands it wrote itself.

How much VRAM do I need?

Enough for a 32B-class model plus a long context — in practice 24 GB and up. The context, not the weights, is what surprises people.

OpenHands or an editor assistant?

An assistant for changes you can describe precisely; OpenHands for multi-step tasks with a clear success criterion, such as making a failing test pass.

Why does the agent keep repeating the same step?

Usually a model that cannot produce the exact tool-call format, or a context window that truncated the plan. Raise the context, then use a larger model.

Recommended hardware

A current option for local AI: GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395). Match memory to your model and software. A mini PC is a complete PC alternative; Mac/MLX and CUDA instructions require compatible hardware.

Amazon Check GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395) price →

As an Amazon Associate, BestLLMfor earns from qualifying purchases, at no extra cost to you. It does not influence our independent rankings.

Did this guide help?

Found an error or have feedback? Let us know — it helps everyone who reads this guide.