Free local copilot: Cline, Tabby & CodeGeeX in VS Code (2026)
GitHub Copilot costs €10/month and sends every character you type to a Microsoft server. Three free VS Code extensions solve both problems by connecting chat and autocompletion to a local LLM: Cline, Tabby, and CodeGeeX. This guide installs all three, shows which models to use, and compares what they actually do to help you choose a free local copilot for VS Code.
#Why use a free local Copilot for VS Code?
Copilot and Cursor send your code to OpenAI or Anthropic. On an NDA-covered client project, with proprietary code, or simply as a matter of principle, that isn’t always acceptable. A local copilot solves three problems at once: zero leakage (nothing leaves the laptop), zero subscription fees, and zero network latency.
The tradeoff is straightforward: the quality of a local Qwen 3.5 9B does not match the very latest Copilot Pro model. But for everyday tasks—completing a loop, writing a test, refactoring a 30-line function—a 9B to 24B model in Q4 gets the job done. And for tasks the local model cannot handle, you can keep Copilot running alongside it.
#Prerequisites
This guide gets you to the model. The kit gets you to the coding copilot in your editor.
- Lifetime online access
- PDF + files
- Lifetime updates
- Recent VS Code
- Version 1.85+ for the inline completion API. JetBrains extensions also exist for all three tools.
- Ollama installed
- Serves as the inference backend for Cline. Listens by default on http://localhost:11434. For Tabby, the runtime is included; CodeGeeX can point to Ollama.
- GPU with 8 GB VRAM minimum
- RTX 3060 12 GB, 4060 Ti 16 GB, unified-memory Mac M2/M3 16 GB. Below that, stick to 1.5B models (qwen2.5-coder:1.5b) for completion; they remain responsive.
- 20 GB of disk space
- A 7B coder in Q4 = 4-5 GB. Allow plenty of headroom if you're testing several models or tools.
#1. Cline (chat + agent)
Cline is the most complete option for assistance: a sidebar chat and agent mode capable of reading and modifying multiple files, Cursor-style, all locally through Ollama. Open source under the MIT license and BYO-LLM, it is the natural replacement for Continue.dev. Roo Code, a fork of Cline, is also shutting down: Cline therefore becomes the consolidation point for this ecosystem. Its limitation: it does not provide inline autocompletion — Tabby covers that lower down.
- 01Install the extensionIn VS Code, open the extensions palette (Ctrl+Shift+X), search for Cline, and install it. A Cline icon appears in the left sidebar.
- 02Download a modelQwen 3.5 9B handles chat and agent tasks respectably on 8 GB of VRAM. Move up to Devstral 24B or Qwen 3.6 35B-A3B if you have the headroom; Devstral is specifically designed for agentic workloads.
- 03Configure the Ollama providerOpen Cline's settings, choose the Ollama provider, enter the local URL, and select the model.
- 04Test the chatAsk a question, paste code, or request a fix. The response appears in the panel without sending anything to the cloud.
- 05Test the agentDescribe a multi-file task. Cline proposes a plan, then applies the changes as diffs that you approve one by one.
#2. Tabby (self-hosted autocomplete)
Tabby (TabbyML) covers exactly what Cline does not: inline autocompletion. It is designed for teams—a single server hosts the model, and all developers connect to it from their IDE. The server includes its own runtime, so it does not need Ollama.
- 01Launch the Tabby serverThe Docker method is the simplest. The server exposes an HTTP endpoint on :8080 and a web admin interface.
- 02Create an admin accountOpen http://localhost:8080 in your browser. On first connection, create the owner account. Go to Settings → Information to retrieve an auth token.
- 03Install the VS Code extensionIn the extensions palette, search for Tabby (publisher: TabbyML). Install it.
- 04Connect the extensionOpen Settings (Ctrl+,), search for tabby. Set Endpoint = http://localhost:8080 and paste the token. The VS Code status bar shows Tabby: ready; completions appear grayed out as you type.
#3. CodeGeeX
CodeGeeX (Tsinghua KEG / Zhipu AI) is the “all-in-one” option: the extension already includes completion and chat with a free proprietary model hosted by them. But it can also connect to a local model, which is what interests us here.
- 01Install the extensionIn VS Code, search for CodeGeeX (publisher: aminer). Install it. A CodeGeeX icon appears in the sidebar.
- 02Switch to local modeOpen the CodeGeeX settings (the gear icon in the side panel). Find the Local mode section and enable it.
- 03Point to OllamaEnter the local model URL. CodeGeeX speaks an OpenAI-compatible endpoint—the one from Ollama on http://localhost:11434/v1 will work.
- 04Choose the modelcodegeex4 (9B) is the built-in model available in Ollama. Otherwise, any recent coder will do (qwen3-coder, devstral).
#Comparison table
| Tool | What it does | Local backend | License | Ideal for |
|---|---|---|---|---|
| Cline | Chat + multi-file agent | Ollama (1 per workstation) | MIT, open source | Solo, power user |
| Tabby | Inline autocomplete | Runtime included (server) | Self-hosted | Team, Cline complement |
| CodeGeeX | Completion + chat | OpenAI-compatible endpoint | Free (cloud by default) | Light all-in-one setup |
- Use-case coverage
- Cline: chat + agent (no completion). Tabby: completion (no advanced chat). CodeGeeX: completion + basic chat. The Cline + Tabby combo covers everything.
- Setup effort
- CodeGeeX (5 min) < Cline + Ollama (10 min) < Tabby self-hosted (15–20 min).
- Multi-utilisateurs
- Only Tabby is designed for this natively. Cline/CodeGeeX = one workstation = one Ollama.
- Privacy
- Cline and Tabby remain local by default. CodeGeeX goes to the cloud until local mode is enabled—keep an eye on that.
#Troubleshooting
- Cline is not responding
- Check ollama ps: the model should appear there when you submit a request. If Ollama does not respond, test curl http://localhost:11434/api/tags from the terminal. Also check that the provider selected in Cline is Ollama.
- No grayed-out completion
- Cline doesn't provide inline autocompletion: that's normal. Install Tabby (or Twinny) for line-by-line completion, and check editor.inlineSuggest.enabled in VS Code.
- Tabby displays “Connection refused”
- The container isn't running, or the port isn't exposed. docker ps should list tabby. docker logs tabby shows the cause (often: the model wasn't downloaded on first launch; wait 5-10 min).
- CodeGeeX returns answers in Chinese
- Specify the language in the system prompt (Settings CodeGeeX → Custom system prompt → "Toujours répondre en français"), or use Qwen 3.5 9B, which responds in French when addressed in French.
- VRAM saturated when everything is running
- You may have two models loaded (Tabby autocompletion + Cline chat). On 8 GB, keep a lightweight FIM model for completion and a Qwen 3.5 9B for chat. ollama ps displays usage.
#FAQ
What’s the best free local alternative to Copilot?+
Why does this guide no longer mention Continue.dev?+
Does Cline provide autocompletion like Copilot?+
Which local model should you choose for coding?+
Do you need a powerful GPU to use it?+
#Go further
You have a local copilot up and running. Three logical directions depending on your profile:
- Learn more about Cline
- The dedicated guide to local Copilot with Cline goes deeper into the configuration: provider Ollama, agent mode, autocompletion via Tabby/Twinny, and fine-tuning the settings for agentic workflows.
- You came from Continue.dev
- The Continue → Cline migration guide details the transition (equivalent settings, models, and local data).
- Choose your hardware
- Before moving up to a 14B or 32B model, the Choose Your GPU for Local AI guide helps you avoid buying too little VRAM.
If you want to save time on configuration, the paid “Local Code Copilot” guide provides a complete Ollama + Cline + Aider configuration pack, ready to paste, for both a solo workstation and a team.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.