Local Copilot with Cline: AI agent in VS Code (Ollama)
GitHub Copilot sends your code to Microsoft. For a freelance developer who signs NDAs, or a team in a regulated industry, that’s a deal-breaker. Cline is the VS Code/JetBrains extension that replaces Copilot with a local model (via Ollama), with a chat and an agent mode capable of reading and modifying multiple files. This guide installs it, configures a recent coding model (Qwen 3.5, Devstral), and shows how to recover the essentials of the Copilot UX—while keeping your code at home.
#Why a local Copilot
- Privacy
- No code snippet leaves your computer. NDA, trade secrets, client code: everything stays local.
- Offline
- On a plane, deep in a lab without Internet, on vacation. The model is on your drive.
- Cost
- Free after buying the GPU. A developer who uses Copilot for two years covers the cost of a mid-range graphics card.
- Model selection
- You can swap Qwen 3.5 for Devstral, Qwen3-Coder, or GLM 4.7 Flash depending on the language and your VRAM.
#Why this guide switched to Cline
This guide originally relied on Continue.dev. Continue was acquired by Cursor in June 2026, its repository is now archived, and the standalone product is no longer maintained, so we've updated it for Cline, an equivalent, active open-source assistant that is also 100% local through Ollama.
This guide gets you to the model. The kit gets you to the coding copilot in your editor.
- Lifetime online access
- PDF + files
- Lifetime updates
#Models to use
- ⭐ Qwen 3.5 9B
- THE 8 GB choice in 2026. Multilingual (including FR), 256k context, vision, Apache 2.0. Good for chat and for Python, TypeScript, and Rust agents. ~6.6 GB in Q4.
- Devstral 24B (Mistral AI)
- Specialist in code agents, Apache 2.0. THE 16 GB choice for Cline's multi-file edits. ~14 GB in Q4.
- Qwen3-Coder 30B-A3B
- Code MoE (3B active, so fast), 256k context. For 24 GB: high-end chat + agent that rivals Copilot. ~19 GB in Q4.
- Qwen 3.8 27B
- The “Copilot-like” model of 2026 (released on 14/08/2026), with 262k context, vision, and Apache 2.0. For 24 GB—switch to “low” reasoning so it does not overthink.
| VRAM | Cline model (chat + agent) | Quantization |
|---|---|---|
| 6 GB | granite4.2:8b | Q4_K_M (~5.3 GB) |
| 8 GB | qwen3.5:9b | Q4_K_M (~6.6 GB) |
| 16 GB | devstral:24b / gpt-oss:20b | Q4 / MXFP4 (~14 GB) |
| 24 GB+ | qwen3-coder:30b / qwen3.8:27b | Q4 (~18-19 GB) |
#1. Install Cline
- 01Install OllamaIf you haven't already, see the guide Install Ollama on Windows/macOS/Linux. This is the engine that runs the model in the background.
- 02Download the modelGet a coder model suited to your VRAM with the command below.
- 03Install the Cline extensionIn VS Code: Extensions (Ctrl+Shift+X) → search for Cline → Install. In JetBrains, use the plugin marketplace.
- 04Open the panelA Cline icon appears in the left sidebar. Click it to open the chat and agent panel.
#2. Configure the Ollama provider
Cline needs to know where to find the model. Open its settings (the gear icon in the panel), select the Ollama provider, enter the daemon’s local address, and choose the model from the detected list.
#3. Chat mode
Chat mode is the direct equivalent of a local Copilot Chat panel. Ask a question, paste code, request an explanation, or ask for a fix.
- Open chat
- Click the Cline icon in the sidebar. The conversation appears in the panel.
- Attach code
- Select a block in the editor, then send it to the chat context so the model can reason about it.
- Mention files
- Cline can reference files in your project to answer with the right context, without requiring you to copy and paste everything.
- Keep the history local
- Your conversations stay on your machine. Nothing is synced to the cloud.
#4. Agent mode
This is where Cline goes beyond a simple chat. In agent mode, it can read the directory tree, open files, suggest multi-file changes, and run commands—each action requires your approval.
- Plan, then execution
- Describe a task in French (“add a /health endpoint and its test”). Cline proposes a plan, then applies the changes step by step.
- Diff validated
- Each change is presented as a diff that you accept or reject. You remain in control at all times.
- Commands in boxes
- The agent may suggest running a command (tests, build). Nothing runs without your explicit approval.
#5. Autocomplete: Tabby or Twinny
Important point: Cline is a chat assistant plus agent; it does not provide inline “grayed-out text” autocomplete while you type. To regain a Copilot-like experience, add a dedicated extension that also connects to Ollama.
- Tabby
- Self-hosted autocompletion extension, ideal if several developers share the same GPU. Line-by-line completion appears in gray; press Tab to accept.
- Twinny
- Lightweight VS Code extension, autocompletion + chat, connecting directly to Ollama. A good solo complement to Cline for high-volume typing.
- The pair that works
- Cline for chat and agent use, Tabby or Twinny for inline completion. Both point to the same Ollama, and your code remains local on both sides.
#Performance and advanced settings
- High num_ctx
- In Ollama, set num_ctx to 16384 so the agent can send a lot of context (files, diffs). Via a Modelfile.
- Flash Attention
- On RTX 30/40/50 cards, enable FA in Ollama (OLLAMA_FLASH_ATTENTION=1 variable). About +20% speed.
- Quantized KV cache
- OLLAMA_KV_CACHE_TYPE=q8_0 halves the context's VRAM usage. Crucial when the agent works with large files.
- One model per role
- A dedicated FIM model (qwen2.5-coder:7b-base, ~4.7 GB) for autocompletion (Tabby/Twinny), a 24B/30B model (Devstral, Qwen3-Coder) for chat and the Cline agent. Ollama loads the one requested.
#Frequently asked questions
Does Cline really replace GitHub Copilot?+
Why not stick with Continue.dev?+
Which model should I choose if I only have 8 GB of VRAM?+
Does my code go to the Internet with Cline?+
Does Cline work on JetBrains, not just VS Code?+
#Conclusion
In about twenty minutes, you have a 100% local code copilot: Cline for chat and agent mode, a recent code model (Qwen 3.5 9B, or Devstral 24B for the agent) running through Ollama as the engine, and Tabby or Twinny (with qwen2.5-coder:7b-base) for inline autocompletion. Your code never leaves your machine, you work offline, and you remain free to switch models based on your languages. If you want to save hours of configuration and start from a proven foundation, the paid “Local Code Copilot” guide provides a complete Ollama + Cline + Aider configuration pack, ready to use.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.