Intermediate 20 minDev

Local Copilot with Cline: AI agent in VS Code (Ollama)

GitHub Copilot sends your code to Microsoft. For a freelance developer who signs NDAs, or a team in a regulated industry, that’s a deal-breaker. Cline is the VS Code/JetBrains extension that replaces Copilot with a local model (via Ollama), with a chat and an agent mode capable of reading and modifying multiple files. This guide installs it, configures a recent coding model (Qwen 3.5, Devstral), and shows how to recover the essentials of the Copilot UX—while keeping your code at home.

By Mohamed Meguedmi·Update 2026-08-27·Tested on Windows, macOS, and Linux
i
Continue is no longer maintained (June 2026)
Continue was acquired by Cursor (Anysphere) on June 18, 2026: the repository is archived (v2.0.0 is the latest version), and the extension is no longer maintained—it remains installable through community forks, and this guide still works as written. For a new setup, we recommend Cline instead (open source, MIT, compatible with Ollama): see our “Migrating from Continue to Cline” guide.

#Why a local Copilot

Privacy
No code snippet leaves your computer. NDA, trade secrets, client code: everything stays local.
Offline
On a plane, deep in a lab without Internet, on vacation. The model is on your drive.
Cost
Free after buying the GPU. A developer who uses Copilot for two years covers the cost of a mid-range graphics card.
Model selection
You can swap Qwen 3.5 for Devstral, Qwen3-Coder, or GLM 4.7 Flash depending on the language and your VRAM.

#Why this guide switched to Cline

This guide originally relied on Continue.dev. Continue was acquired by Cursor in June 2026, its repository is now archived, and the standalone product is no longer maintained, so we've updated it for Cline, an equivalent, active open-source assistant that is also 100% local through Ollama.

The Local Copilot Kit

This guide gets you to the model. The kit gets you to the coding copilot in your editor.

  • Lifetime online access
  • PDF + files
  • Lifetime updates
i
Were you coming from Continue?
The transition is painless: your Ollama backend does not change; you only replace the extension. Your pinned Continue forks remain legally runnable (Apache 2.0 code), but note that the export window for hosted data closed on July 15, 2026. A dedicated guide provides a step-by-step migration if needed.

#Models to use

⭐ Qwen 3.5 9B
THE 8 GB choice in 2026. Multilingual (including FR), 256k context, vision, Apache 2.0. Good for chat and for Python, TypeScript, and Rust agents. ~6.6 GB in Q4.
Devstral 24B (Mistral AI)
Specialist in code agents, Apache 2.0. THE 16 GB choice for Cline's multi-file edits. ~14 GB in Q4.
Qwen3-Coder 30B-A3B
Code MoE (3B active, so fast), 256k context. For 24 GB: high-end chat + agent that rivals Copilot. ~19 GB in Q4.
Qwen 3.8 27B
The “Copilot-like” model of 2026 (released on 14/08/2026), with 262k context, vision, and Apache 2.0. For 24 GB—switch to “low” reasoning so it does not overthink.
Recommended model based on available VRAM
VRAMCline model (chat + agent)Quantization
6 GBgranite4.2:8bQ4_K_M (~5.3 GB)
8 GBqwen3.5:9bQ4_K_M (~6.6 GB)
16 GBdevstral:24b / gpt-oss:20bQ4 / MXFP4 (~14 GB)
24 GB+qwen3-coder:30b / qwen3.8:27bQ4 (~18-19 GB)
→
Cline prefers a model that follows instructions
Unlike raw autocomplete, agent mode sends multi-step instructions. Prefer the model's instruct/chat variant (not the -base variant), and move up to a larger model (Devstral 24B, Qwen3-Coder 30B) if your VRAM allows for greater reliability.

#1. Install Cline

  1. 01
    Install Ollama
    If you haven't already, see the guide Install Ollama on Windows/macOS/Linux. This is the engine that runs the model in the background.
  2. 02
    Download the model
    Get a coder model suited to your VRAM with the command below.
  3. 03
    Install the Cline extension
    In VS Code: Extensions (Ctrl+Shift+X) → search for Cline → Install. In JetBrains, use the plugin marketplace.
  4. 04
    Open the panel
    A Cline icon appears in the left sidebar. Click it to open the chat and agent panel.
Model to download
# Modèle principal (chat + agent)
ollama pull qwen3.5:9b

# Plus de VRAM ? Devstral, le spécialiste du mode agent (16 Go)
ollama pull devstral:24b

#2. Configure the Ollama provider

Cline needs to know where to find the model. Open its settings (the gear icon in the panel), select the Ollama provider, enter the daemon’s local address, and choose the model from the detected list.

Provider settings
Provider     : Ollama
Base URL     : http://localhost:11434
Modèle       : qwen3.5:9b
!
Stay firmly with the local provider
Cline can also talk to cloud APIs. For 100% local use, verify that the active provider is Ollama and not a remote provider—that is the only setting separating private use from sending code to a third-party server.

#3. Chat mode

Chat mode is the direct equivalent of a local Copilot Chat panel. Ask a question, paste code, request an explanation, or ask for a fix.

Open chat
Click the Cline icon in the sidebar. The conversation appears in the panel.
Attach code
Select a block in the editor, then send it to the chat context so the model can reason about it.
Mention files
Cline can reference files in your project to answer with the right context, without requiring you to copy and paste everything.
Keep the history local
Your conversations stay on your machine. Nothing is synced to the cloud.

#4. Agent mode

This is where Cline goes beyond a simple chat. In agent mode, it can read the directory tree, open files, suggest multi-file changes, and run commands—each action requires your approval.

Plan, then execution
Describe a task in French (“add a /health endpoint and its test”). Cline proposes a plan, then applies the changes step by step.
Diff validated
Each change is presented as a diff that you accept or reject. You remain in control at all times.
Commands in boxes
The agent may suggest running a command (tests, build). Nothing runs without your explicit approval.
Examples of instructions for the agent
- Ajoute du logging pour tracer les entrées/sorties
- Convertis cette fonction en async/await
- Extrais ce bloc dans un fichier util.ts et mets à jour les imports
- Écris des tests unitaires pour les cas limites
→
Give the agent context
The more precise the instruction (target file, expected behavior, style constraint), the better the result. A 24B model (Devstral) follows multi-step instructions more faithfully than a 9B model.

#5. Autocomplete: Tabby or Twinny

Important point: Cline is a chat assistant plus agent; it does not provide inline “grayed-out text” autocomplete while you type. To regain a Copilot-like experience, add a dedicated extension that also connects to Ollama.

Tabby
Self-hosted autocompletion extension, ideal if several developers share the same GPU. Line-by-line completion appears in gray; press Tab to accept.
Twinny
Lightweight VS Code extension, autocompletion + chat, connecting directly to Ollama. A good solo complement to Cline for high-volume typing.
The pair that works
Cline for chat and agent use, Tabby or Twinny for inline completion. Both point to the same Ollama, and your code remains local on both sides.
→
Dedicated completion model
For autocomplete (Fill-In-the-Middle), qwen2.5-coder:7b-base remains the benchmark in 2026 (~4.7 GB) and is still highly responsive—it is one of the few cases where a 2024 model remains at the top. Keep your chat/agent model (Qwen 3.5 9B, Devstral 24B) for Cline.

#Performance and advanced settings

High num_ctx
In Ollama, set num_ctx to 16384 so the agent can send a lot of context (files, diffs). Via a Modelfile.
Flash Attention
On RTX 30/40/50 cards, enable FA in Ollama (OLLAMA_FLASH_ATTENTION=1 variable). About +20% speed.
Quantized KV cache
OLLAMA_KV_CACHE_TYPE=q8_0 halves the context's VRAM usage. Crucial when the agent works with large files.
One model per role
A dedicated FIM model (qwen2.5-coder:7b-base, ~4.7 GB) for autocompletion (Tabby/Twinny), a 24B/30B model (Devstral, Qwen3-Coder) for chat and the Cline agent. Ollama loads the one requested.
Increase the context window
# Modelfile : devstral-16k
FROM devstral:24b
PARAMETER num_ctx 16384
Create the extended-context model
ollama create devstral-16k -f Modelfile
# puis sélectionnez devstral-16k comme modèle dans Cline

#Frequently asked questions

FAQ
Does Cline really replace GitHub Copilot?+
For chat and agent mode (reading/modifying multiple files, running commands), yes. For inline “grayed-out text” autocompletion while typing, no: Cline does not provide it. Add Tabby or Twinny, connected to the same Ollama, to get that convenience back.
Why not stick with Continue.dev?+
Continue was acquired by Cursor in June 2026; development continues at Cursor (v2.0.0 = latest stable release), but the future of the standalone product is uncertain. The extension can still be installed, and the Apache 2.0 code remains legal to run, but there are no updates or fixes. Cline, MIT-licensed, open source, and actively developed, is the recommended consolidation choice.
Which model should I choose if I only have 8 GB of VRAM?+
Qwen 3.5 9B in Q4_K_M (~6,6 GB) fits comfortably on 8 GB and remains the best choice in this range in 2026 (256k context, multilingual, vision). If VRAM is really tight, Granite 4.2 8B (~5,3 GB) is even more modest. Both are sufficient for chat; for a more reliable multi-step agent mode, move up to Devstral 24B as soon as your VRAM (16 GB) allows.
Does my code go to the Internet with Cline?+
No, as long as the active provider is Ollama (http://localhost:11434). Cline can also connect to cloud APIs: check its settings to make sure you are using the local Ollama provider; this is the only setting that separates 100% private use from sending data to a third-party server.
Does Cline work on JetBrains, not just VS Code?+
Yes. Cline is available for VS Code and JetBrains IDEs (through the plugin marketplace). The Ollama provider configuration and model selection are identical on both.

#Conclusion

In about twenty minutes, you have a 100% local code copilot: Cline for chat and agent mode, a recent code model (Qwen 3.5 9B, or Devstral 24B for the agent) running through Ollama as the engine, and Tabby or Twinny (with qwen2.5-coder:7b-base) for inline autocompletion. Your code never leaves your machine, you work offline, and you remain free to switch models based on your languages. If you want to save hours of configuration and start from a proven foundation, the paid “Local Code Copilot” guide provides a complete Ollama + Cline + Aider configuration pack, ready to use.

Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.