Intermediate 15 minIDE

Free local copilot: Cline, Tabby & CodeGeeX in VS Code (2026)

GitHub Copilot costs €10/month and sends every character you type to a Microsoft server. Three free VS Code extensions solve both problems by connecting chat and autocompletion to a local LLM: Cline, Tabby, and CodeGeeX. This guide installs all three, shows which models to use, and compares what they actually do to help you choose a free local copilot for VS Code.

By Mohamed Meguedmi·Update 2026-08-27·Tested on Windows, macOS, and Linux
i
Continue is no longer maintained (June 2026)
Continue was acquired by Cursor (Anysphere) on June 18, 2026: the repository is archived (v2.0.0 is the latest version), and the extension is no longer maintained—it remains installable through community forks, and this guide still works as written. For a new setup, we recommend Cline instead (open source, MIT, compatible with Ollama): see our “Migrating from Continue to Cline” guide.

#Why use a free local Copilot for VS Code?

Copilot and Cursor send your code to OpenAI or Anthropic. On an NDA-covered client project, with proprietary code, or simply as a matter of principle, that isn’t always acceptable. A local copilot solves three problems at once: zero leakage (nothing leaves the laptop), zero subscription fees, and zero network latency.

The tradeoff is straightforward: the quality of a local Qwen 3.5 9B does not match the very latest Copilot Pro model. But for everyday tasks—completing a loop, writing a test, refactoring a 30-line function—a 9B to 24B model in Q4 gets the job done. And for tasks the local model cannot handle, you can keep Copilot running alongside it.

i
What these 3 tools do
Cline = chat + multi-file agent (no inline autocompletion). Tabby = autocompletion only, focused on self-hosted teams. CodeGeeX = completion + chat, with its own integrated model. All work in VS Code and JetBrains.
→
What about Continue.dev?
Continue.dev appeared in roundups through 2026, but it was acquired by Cursor (announced on June 18, 2026), and v2.0.0 is the latest stable release published; the future of the standalone extension is uncertain. We replaced it with Cline, its active open-source equivalent. If you were still using it, note that the export window for hosted data closed on July 15, 2026.

#Prerequisites

The Local Copilot Kit

This guide gets you to the model. The kit gets you to the coding copilot in your editor.

  • Lifetime online access
  • PDF + files
  • Lifetime updates
Recent VS Code
Version 1.85+ for the inline completion API. JetBrains extensions also exist for all three tools.
Ollama installed
Serves as the inference backend for Cline. Listens by default on http://localhost:11434. For Tabby, the runtime is included; CodeGeeX can point to Ollama.
GPU with 8 GB VRAM minimum
RTX 3060 12 GB, 4060 Ti 16 GB, unified-memory Mac M2/M3 16 GB. Below that, stick to 1.5B models (qwen2.5-coder:1.5b) for completion; they remain responsive.
20 GB of disk space
A 7B coder in Q4 = 4-5 GB. Allow plenty of headroom if you're testing several models or tools.
→
Actual VRAM by coder size
Qwen 3.5 9B Q4 ≈ 6.6 GB · Gemma 4 12B ≈ 7.6 GB · Devstral 24B Q4 ≈ 14 GB · Qwen 3.6 35B-A3B ≈ 23 GB. For inline completion (FIM), latency matters more than quality: a small FIM model such as Qwen2.5-Coder is sufficient. Keep a Qwen 3.5 9B for chat and a Devstral 24B for the agent if you have the VRAM.

#1. Cline (chat + agent)

Cline is the most complete option for assistance: a sidebar chat and agent mode capable of reading and modifying multiple files, Cursor-style, all locally through Ollama. Open source under the MIT license and BYO-LLM, it is the natural replacement for Continue.dev. Roo Code, a fork of Cline, is also shutting down: Cline therefore becomes the consolidation point for this ecosystem. Its limitation: it does not provide inline autocompletion — Tabby covers that lower down.

  1. 01
    Install the extension
    In VS Code, open the extensions palette (Ctrl+Shift+X), search for Cline, and install it. A Cline icon appears in the left sidebar.
  2. 02
    Download a model
    Qwen 3.5 9B handles chat and agent tasks respectably on 8 GB of VRAM. Move up to Devstral 24B or Qwen 3.6 35B-A3B if you have the headroom; Devstral is specifically designed for agentic workloads.
  3. 03
    Configure the Ollama provider
    Open Cline's settings, choose the Ollama provider, enter the local URL, and select the model.
  4. 04
    Test the chat
    Ask a question, paste code, or request a fix. The response appears in the panel without sending anything to the cloud.
  5. 05
    Test the agent
    Describe a multi-file task. Cline proposes a plan, then applies the changes as diffs that you approve one by one.
Download the model
ollama pull qwen3.5:9b
# Plus de VRAM ? Devstral, spécialiste agent de code, suit mieux le mode agent
ollama pull devstral:24b
Cline provider settings
Provider : Ollama
Base URL : http://localhost:11434
Modèle   : qwen3.5:9b
→
Instruct variant for the agent
For Cline's chat and agent, use the model's instruct/chat variant (not the -base variant, which is reserved for autocompletion). It follows multi-step instructions better.

#2. Tabby (self-hosted autocomplete)

Tabby (TabbyML) covers exactly what Cline does not: inline autocompletion. It is designed for teams—a single server hosts the model, and all developers connect to it from their IDE. The server includes its own runtime, so it does not need Ollama.

  1. 01
    Launch the Tabby server
    The Docker method is the simplest. The server exposes an HTTP endpoint on :8080 and a web admin interface.
  2. 02
    Create an admin account
    Open http://localhost:8080 in your browser. On first connection, create the owner account. Go to Settings → Information to retrieve an auth token.
  3. 03
    Install the VS Code extension
    In the extensions palette, search for Tabby (publisher: TabbyML). Install it.
  4. 04
    Connect the extension
    Open Settings (Ctrl+,), search for tabby. Set Endpoint = http://localhost:8080 and paste the token. The VS Code status bar shows Tabby: ready; completions appear grayed out as you type.
Tabby startup (GPU NVIDIA)
docker run -d --name tabby \
  --gpus all -p 8080:8080 \
  -v $HOME/.tabby:/data \
  registry.tabbyml.com/tabbyml/tabby \
  serve --model StarCoder-1B --device cuda
i
Cline + Tabby: the complete duo
Cline handles chat and the agent, while Tabby handles inline autocompletion. Together, they recreate the full Copilot experience — locally and for free. This is the recommended combination for both a solo workstation and a team. Twinny is an alternative to Tabby if you prefer to stay on Ollama for completion as well.

#3. CodeGeeX

CodeGeeX (Tsinghua KEG / Zhipu AI) is the “all-in-one” option: the extension already includes completion and chat with a free proprietary model hosted by them. But it can also connect to a local model, which is what interests us here.

  1. 01
    Install the extension
    In VS Code, search for CodeGeeX (publisher: aminer). Install it. A CodeGeeX icon appears in the sidebar.
  2. 02
    Switch to local mode
    Open the CodeGeeX settings (the gear icon in the side panel). Find the Local mode section and enable it.
  3. 03
    Point to Ollama
    Enter the local model URL. CodeGeeX speaks an OpenAI-compatible endpoint—the one from Ollama on http://localhost:11434/v1 will work.
  4. 04
    Choose the model
    codegeex4 (9B) is the built-in model available in Ollama. Otherwise, any recent coder will do (qwen3-coder, devstral).
CodeGeeX model in Ollama
ollama pull codegeex4
!
Default cloud mode
By default, CodeGeeX uses Zhipu's servers in China. Unless you explicitly switch to local mode in the settings, your code is sent to them. Check the badge in the VS Code status bar: it should display Local, not Cloud.

#Comparison table

The 3 free local alternatives to Copilot in VS Code
ToolWhat it doesLocal backendLicenseIdeal for
ClineChat + multi-file agentOllama (1 per workstation)MIT, open sourceSolo, power user
TabbyInline autocompleteRuntime included (server)Self-hostedTeam, Cline complement
CodeGeeXCompletion + chatOpenAI-compatible endpointFree (cloud by default)Light all-in-one setup
Use-case coverage
Cline: chat + agent (no completion). Tabby: completion (no advanced chat). CodeGeeX: completion + basic chat. The Cline + Tabby combo covers everything.
Setup effort
CodeGeeX (5 min) < Cline + Ollama (10 min) < Tabby self-hosted (15–20 min).
Multi-utilisateurs
Only Tabby is designed for this natively. Cline/CodeGeeX = one workstation = one Ollama.
Privacy
Cline and Tabby remain local by default. CodeGeeX goes to the cloud until local mode is enabled—keep an eye on that.
→
Practical recommendation
To get started on your own: Cline + Qwen 3.5 9B on Ollama for chat and agent work, plus Tabby for autocompletion. For a team of 3+ devs sharing a GPU: Tabby for shared completion, Cline on the workstation. CodeGeeX remains a lightweight all-in-one alternative.

#Troubleshooting

Cline is not responding
Check ollama ps: the model should appear there when you submit a request. If Ollama does not respond, test curl http://localhost:11434/api/tags from the terminal. Also check that the provider selected in Cline is Ollama.
No grayed-out completion
Cline doesn't provide inline autocompletion: that's normal. Install Tabby (or Twinny) for line-by-line completion, and check editor.inlineSuggest.enabled in VS Code.
Tabby displays “Connection refused”
The container isn't running, or the port isn't exposed. docker ps should list tabby. docker logs tabby shows the cause (often: the model wasn't downloaded on first launch; wait 5-10 min).
CodeGeeX returns answers in Chinese
Specify the language in the system prompt (Settings CodeGeeX → Custom system prompt → "Toujours répondre en français"), or use Qwen 3.5 9B, which responds in French when addressed in French.
VRAM saturated when everything is running
You may have two models loaded (Tabby autocompletion + Cline chat). On 8 GB, keep a lightweight FIM model for completion and a Qwen 3.5 9B for chat. ollama ps displays usage.

#FAQ

What’s the best free local alternative to Copilot?+
There isn't a single answer because the tools don't address the same need. For chat and multi-file agents, Cline is the most complete. For inline autocompletion, Tabby (or Twinny). In practice, the Cline + Tabby duo recreates the full Copilot experience for free and locally. CodeGeeX is the fastest all-in-one option to install.
Why does this guide no longer mention Continue.dev?+
Continue.dev was acquired by Cursor (acqui-hire announced on June 18, 2026). v2.0.0 is the latest stable version released, and the future of the standalone extension is uncertain. The code remains under Apache 2.0, so pinned builds continue to work, but for a living codebase, migrating to Cline is preferable. The deadline for exporting hosted data was July 15, 2026 (now past).
Does Cline provide autocompletion like Copilot?+
No. Cline focuses on chat and agent mode (reading and modifying multiple files). It does not offer line-by-line gray-text completion. For that feature, add Tabby or Twinny in parallel—the two extensions coexist without conflicts in VS Code.
Which local model should you choose for coding?+
For inline completion (FIM), Qwen2.5-Coder remains a safe choice thanks to its low latency. For chat and agents: Qwen 3.5 9B on 8 GB, Devstral 24B or gpt-oss 20B on 16 GB, Qwen 3.6 35B-A3B if you have the headroom. Devstral, a code-agent specialist, is an excellent alternative. Simple rule: use a lightweight model for completion and a larger model for reasoning.
Do you need a powerful GPU to use it?+
No. 8 GB of VRAM is enough for a Qwen 3.5 9B in Q4 (chat) plus a small FIM model for completion. Below that, stick to lightweight FIM models that remain responsive. A Mac M2/M3 with 16 GB of unified memory also handles the job very well.

#Go further

You have a local copilot up and running. Three logical directions depending on your profile:

Learn more about Cline
The dedicated guide to local Copilot with Cline goes deeper into the configuration: provider Ollama, agent mode, autocompletion via Tabby/Twinny, and fine-tuning the settings for agentic workflows.
You came from Continue.dev
The Continue → Cline migration guide details the transition (equivalent settings, models, and local data).
Choose your hardware
Before moving up to a 14B or 32B model, the Choose Your GPU for Local AI guide helps you avoid buying too little VRAM.

If you want to save time on configuration, the paid “Local Code Copilot” guide provides a complete Ollama + Cline + Aider configuration pack, ready to paste, for both a solo workstation and a team.

Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.

Prices in euros (€) are French market prices including VAT, checked by QuelLLM. US prices differ: the Amazon buttons show the current US price.