Intermediate 14 minAgents

Cline + Ollama: 100% local coding agent in VS Code

Cline turns VS Code into an autonomous coding agent: it reads your files, proposes diffs, runs commands, and iterates until the task is finished. Connected to Ollama, all this work stays on your machine — no token is sent to a third-party server. This guide walks you through a complete, working Cline + Ollama setup: which model to choose based on your VRAM, the Ollama configuration that avoids unpleasant context surprises, and the settings that make the difference between a usable agent and one that runs in circles.

By Mohamed Meguedmi·Update 2026-08-27·Tested on Windows, macOS, and Linux

#Why run a Cline agent locally

Cline (formerly Claude Dev) is a VS Code extension that drives an LLM in an agent loop: it plans, edits multiple files, runs terminal commands, reads the output, and fixes issues. Unlike simple inline completion, it executes complete tasks—writing a function, migrating a module, or debugging a stack trace.

Pairing it with Ollama instead of a cloud API has three concrete benefits: your proprietary code never leaves the machine, there is no per-token cost (Cline consumes a huge amount of context, which adds up quickly in the cloud), and the agent works offline. The tradeoff: a good local coding model requires VRAM, and Ollama's default settings are not suitable for agent use.

i
This guide is not a migration guide
If you’re coming from Continue.dev and want to recover your config, follow our dedicated guide, “Continue.dev acquired by Cursor: migrate to Cline”—it covers data export and settings mapping. Here, we start from scratch to build a clean Cline + Ollama setup.

#Prerequisites

The Local Copilot Kit

This guide gets you to the model. The kit gets you to the coding copilot in your editor.

  • Lifetime online access
  • PDF + files
  • Lifetime updates
VS Code
Recent version (1.90+). Cline also works in compatible Open VSX forks (VSCodium, Cursor), but this guide targets standard VS Code.
Ollama installed
The daemon must be running and listening on http://localhost:11434. If it isn't already, see our Ollama installation guide.
A GPU with enough VRAM
16 GB minimum for a comfortable agent, 24 GB for top-tier performance. A pure CPU setup remains possible but is too slow for the agent loop.
A project opened in VS Code
Cline operates in the current workspace folder: open a real repository, not an isolated file.

#Choose the model by VRAM

Model choice matters more than anything else: a Cline agent chains tool calls, file reads, and diffs, so it needs a model that follows instructions and respects tool formats. Three profiles cover most setups in 2026.

24 GB (RTX 4090, 3090, RX 7900 XTX) — Qwen3-Coder 32B
The best local agent today. In Q4_K_M, it uses ~19 GB, leaving room for a 32k context. This is the model to target if your card can handle it.
16 GB (RTX 4080, 5070 Ti, 4060 Ti 16GB) — Devstral Small 2
Mistral model specialized for coding agents, designed to fit in 16 GB in Q4_K_M with a useful context window. The best compromise in this card lineup.
Solid general-purpose model — Qwen 3.6 27B
If you want a single model for code and everything else, the 27B variant fits on 24 GB in Q4 and performs very well as an agent. A notch below Qwen3-Coder for pure coding, but more versatile.
→
VRAM benchmark for Q4
With Q4_K_M quantization (the recommended default): a 14B model ≈ 9 GB, a 27B model ≈ 16–17 GB, and a 32B model ≈ 19 GB. Add the context’s VRAM: the larger the window, the larger the KV cache. Always keep 2–3 GB of headroom.

Avoid models below 14B for agent use: they break the tool-call format, get stuck in loops, or propose diffs that do not apply. A small, fast model is excellent for inline completion (via Tabby or Continue), but not for driving Cline.

#1. Prepare Ollama

First verify that Ollama is running, then pull the selected model. We’ll use Qwen3-Coder 32B as an example here—replace the tag with the one that matches your VRAM.

Terminal
# Vérifier que le daemon répond
curl http://localhost:11434/api/version

# Tirer le modèle de code (adapter au tag exact affiché sur ollama.com)
ollama pull qwen3-coder:32b

# Alternatives selon la VRAM
ollama pull devstral-small-2   # 16 Go
ollama pull qwen3.6:27b        # généraliste 24 Go

# Lister ce qui est installé
ollama list
!
Check the exact tag
Ollama tag names evolve (quantization suffixes, versions). Copy the tag from the model's official page on ollama.com rather than guessing it: a nonexistent tag returns a silent pull error.

#2. Adjust the Ollama context

This is the step everyone misses. By default, Ollama serves models with a 4096-token context window. But Cline sends the contents of several files, the task history, and tool definitions: at 4k, the agent forgets the beginning of its own task on every turn and becomes unusable.

num_ctx needs to be increased. The proper way is to create a derived model through a Modelfile, so the setting is permanent and independent of Cline.

Modelfile
# Fichier : Modelfile
FROM qwen3-coder:32b

# Contexte élargi pour l'usage agent (32k)
PARAMETER num_ctx 32768

# Un peu de déterminisme pour le code
PARAMETER temperature 0.2
Terminal
# Créer le modèle dérivé
ollama create qwen3-coder-agent -f Modelfile

# Il apparaît désormais dans la liste
ollama list
!
num_ctx costs VRAM
Going from 4k to 32k increases the KV cache by several GB. On 16 GB, stick to 16k (num_ctx 16384) rather than 32k; otherwise, Ollama spills the model over the available RAM and slows the agent down. Monitor `ollama ps`: if the PROCESSOR column shows CPU, reduce the context or quantization.

#3. Install Cline

  1. 01
    Open the Marketplace
    In VS Code, open the Extensions panel (Ctrl+Shift+X).
  2. 02
    Search for Cline
    Type “Cline” in the search bar. The official extension is published by “Cline Bot Inc.” — check the publisher to avoid clones.
  3. 03
    Install and pin
    Click Install. A Cline icon appears in the activity bar; pin it for quick access.

#4. Connect Cline to Ollama

Cline handles Ollama natively: no need to cobble together an OpenAI-compatible API URL. Open the Cline panel, click the settings icon, and configure the provider.

  1. 01
    API Provider → Ollama
    From the “API Provider” dropdown, select “Ollama”.
  2. 02
    Base URL
    Leave http://localhost:11434 at its default value. Change it only if Ollama runs on another machine or in a container.
  3. 03
    Model
    Select qwen3-coder-agent (the derived model created in step 2), not the raw model with 4k context.
  4. 04
    Check the context window
    Cline displays a “Context Window” field. Set it to match the Modelfile's num_ctx (32768). A mismatch here causes prompts to be truncated on the Cline side before Ollama.
i
Plan Mode and Act Mode
Cline distinguishes Plan mode (it thinks and proposes a strategy without touching the files) from Act mode (it executes). With a local model, start in Plan mode to validate the approach: this prevents a somewhat weak model from editing ten files in the wrong direction.

#5. First run in agent mode

Open a real project, then give Cline a concrete, well-scoped task. Good first tests touch few files: add a test, fix an identified bug, or write a small utility function.

Cline prompt
Ajoute une fonction `slugify(texte)` dans src/utils/strings.js qui met en minuscules, remplace les accents et les espaces par des tirets. Écris aussi un test unitaire dans le fichier de tests existant à côté.

Cline will read the folder, propose a diff, and ask you to approve each edit and command. Approve step by step the first time to understand its behavior. Once you are comfortable, you can enable auto-approval for certain actions (file reading, nondestructive commands) in the settings.

→
A project rules file
Add a `.clinerules` file to the repository root to set boundaries for the agent: naming conventions, the test command to run, and directories it must never touch. A local model follows instructions better when the boundaries are explicit and brief.

#Common pitfalls

The agent forgets its task
Almost always a context issue: num_ctx is too low on the Ollama side or the Context Window is set incorrectly on the Cline side. Align the two values.
Diffs do not apply
The model is too small or too heavily quantized to preserve the editing format. Move up a tier (14B → 27B/32B) or switch from Q3 to Q4_K_M.
Inference switches to the CPU
The model plus the context exceed the VRAM. Check `ollama ps`; reduce num_ctx or choose a smaller model. An agent running partly on the CPU becomes unusable.
Infinite tool loops
Some models repeat the same command. Switch to Plan Mode first, and break the task into smaller subtasks.
Connection refused
Ollama isn't listening, or not at the expected URL. Test `curl http://localhost:11434/api/version`. In a container, expose the port and point Cline to the correct address.

#Go further

You have a 100% local Cline agent that edits your code without exposing anything. Two natural next steps: fine-tune your model choice and supplement it with the inline autocompletion that Cline doesn't provide.

→
Local Copilot Kit
The complete local replacement for GitHub Copilot: Cline (chat + agent) for tasks, Tabby for gray autocomplete while you type, and a code model suited to your GPU. Each component stays on your machine.
Best local LLM for coding in 2026
The Devstral / Qwen3-Coder comparison and alternatives, with VRAM and code quality, to help refine the model connected to Cline.
Continue.dev acquired by Cursor: migrate to Cline
If you are coming from Continue.dev, see the dedicated guide to exporting and migrating your config to this same setup.
Choose your quantization (Q4, Q5, Q8, FP16)
Understand the quality/VRAM tradeoffs to fit the largest possible model on your card.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.