Cline + Ollama: 100% local coding agent in VS Code
Cline turns VS Code into an autonomous coding agent: it reads your files, proposes diffs, runs commands, and iterates until the task is finished. Connected to Ollama, all this work stays on your machine — no token is sent to a third-party server. This guide walks you through a complete, working Cline + Ollama setup: which model to choose based on your VRAM, the Ollama configuration that avoids unpleasant context surprises, and the settings that make the difference between a usable agent and one that runs in circles.
#Why run a Cline agent locally
Cline (formerly Claude Dev) is a VS Code extension that drives an LLM in an agent loop: it plans, edits multiple files, runs terminal commands, reads the output, and fixes issues. Unlike simple inline completion, it executes complete tasks—writing a function, migrating a module, or debugging a stack trace.
Pairing it with Ollama instead of a cloud API has three concrete benefits: your proprietary code never leaves the machine, there is no per-token cost (Cline consumes a huge amount of context, which adds up quickly in the cloud), and the agent works offline. The tradeoff: a good local coding model requires VRAM, and Ollama's default settings are not suitable for agent use.
#Prerequisites
This guide gets you to the model. The kit gets you to the coding copilot in your editor.
- Lifetime online access
- PDF + files
- Lifetime updates
- VS Code
- Recent version (1.90+). Cline also works in compatible Open VSX forks (VSCodium, Cursor), but this guide targets standard VS Code.
- Ollama installed
- The daemon must be running and listening on http://localhost:11434. If it isn't already, see our Ollama installation guide.
- A GPU with enough VRAM
- 16 GB minimum for a comfortable agent, 24 GB for top-tier performance. A pure CPU setup remains possible but is too slow for the agent loop.
- A project opened in VS Code
- Cline operates in the current workspace folder: open a real repository, not an isolated file.
#Choose the model by VRAM
Model choice matters more than anything else: a Cline agent chains tool calls, file reads, and diffs, so it needs a model that follows instructions and respects tool formats. Three profiles cover most setups in 2026.
- 24 GB (RTX 4090, 3090, RX 7900 XTX) — Qwen3-Coder 32B
- The best local agent today. In Q4_K_M, it uses ~19 GB, leaving room for a 32k context. This is the model to target if your card can handle it.
- 16 GB (RTX 4080, 5070 Ti, 4060 Ti 16GB) — Devstral Small 2
- Mistral model specialized for coding agents, designed to fit in 16 GB in Q4_K_M with a useful context window. The best compromise in this card lineup.
- Solid general-purpose model — Qwen 3.6 27B
- If you want a single model for code and everything else, the 27B variant fits on 24 GB in Q4 and performs very well as an agent. A notch below Qwen3-Coder for pure coding, but more versatile.
Avoid models below 14B for agent use: they break the tool-call format, get stuck in loops, or propose diffs that do not apply. A small, fast model is excellent for inline completion (via Tabby or Continue), but not for driving Cline.
#1. Prepare Ollama
First verify that Ollama is running, then pull the selected model. We’ll use Qwen3-Coder 32B as an example here—replace the tag with the one that matches your VRAM.
#2. Adjust the Ollama context
This is the step everyone misses. By default, Ollama serves models with a 4096-token context window. But Cline sends the contents of several files, the task history, and tool definitions: at 4k, the agent forgets the beginning of its own task on every turn and becomes unusable.
num_ctx needs to be increased. The proper way is to create a derived model through a Modelfile, so the setting is permanent and independent of Cline.
#3. Install Cline
- 01Open the MarketplaceIn VS Code, open the Extensions panel (Ctrl+Shift+X).
- 02Search for ClineType “Cline” in the search bar. The official extension is published by “Cline Bot Inc.” — check the publisher to avoid clones.
- 03Install and pinClick Install. A Cline icon appears in the activity bar; pin it for quick access.
#4. Connect Cline to Ollama
Cline handles Ollama natively: no need to cobble together an OpenAI-compatible API URL. Open the Cline panel, click the settings icon, and configure the provider.
- 01API Provider → OllamaFrom the “API Provider” dropdown, select “Ollama”.
- 02Base URLLeave http://localhost:11434 at its default value. Change it only if Ollama runs on another machine or in a container.
- 03ModelSelect qwen3-coder-agent (the derived model created in step 2), not the raw model with 4k context.
- 04Check the context windowCline displays a “Context Window” field. Set it to match the Modelfile's num_ctx (32768). A mismatch here causes prompts to be truncated on the Cline side before Ollama.
#5. First run in agent mode
Open a real project, then give Cline a concrete, well-scoped task. Good first tests touch few files: add a test, fix an identified bug, or write a small utility function.
Cline will read the folder, propose a diff, and ask you to approve each edit and command. Approve step by step the first time to understand its behavior. Once you are comfortable, you can enable auto-approval for certain actions (file reading, nondestructive commands) in the settings.
#Common pitfalls
- The agent forgets its task
- Almost always a context issue: num_ctx is too low on the Ollama side or the Context Window is set incorrectly on the Cline side. Align the two values.
- Diffs do not apply
- The model is too small or too heavily quantized to preserve the editing format. Move up a tier (14B → 27B/32B) or switch from Q3 to Q4_K_M.
- Inference switches to the CPU
- The model plus the context exceed the VRAM. Check `ollama ps`; reduce num_ctx or choose a smaller model. An agent running partly on the CPU becomes unusable.
- Infinite tool loops
- Some models repeat the same command. Switch to Plan Mode first, and break the task into smaller subtasks.
- Connection refused
- Ollama isn't listening, or not at the expected URL. Test `curl http://localhost:11434/api/version`. In a container, expose the port and point Cline to the correct address.
#Go further
You have a 100% local Cline agent that edits your code without exposing anything. Two natural next steps: fine-tune your model choice and supplement it with the inline autocompletion that Cline doesn't provide.
- Best local LLM for coding in 2026
- The Devstral / Qwen3-Coder comparison and alternatives, with VRAM and code quality, to help refine the model connected to Cline.
- Continue.dev acquired by Cursor: migrate to Cline
- If you are coming from Continue.dev, see the dedicated guide to exporting and migrating your config to this same setup.
- Choose your quantization (Q4, Q5, Q8, FP16)
- Understand the quality/VRAM tradeoffs to fit the largest possible model on your card.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.