Tabby: install local code autocompletion self-hosted
Cline and Aider provide chat and agent mode, but they don’t provide the inline “gray text” autocompletion that appears as you type—the hallmark of GitHub Copilot. Tabby fills that exact gap: it’s a self-hosted, open-source code-completion server (Apache 2.0) that runs in Docker on your GPU and provides Fill-In-the-Middle suggestions directly in the editor. This guide installs it step by step, configures a StarCoder or Qwen-Coder model, generates the token, connects VS Code and JetBrains, and shows how a single GPU can serve an entire team—your code never leaves your infrastructure.
#Why Tabby instead of Cline
Tabby and Cline are not competitors: they complement each other. Cline (and Aider) excels at conversational chat and multi-file agent mode, but neither provides line-by-line autocomplete while you type. Tabby does exactly that—and does it very well. The winning combination in 2026: Cline for discussion and refactoring, Tabby for the gray suggestion that appears as you type.
- Inline grayed-out text
- The suggestion appears directly after the cursor, grayed out. Press Tab to accept, Esc to ignore. Exactly Copilot's UX.
- 100% self-hosted
- The server runs locally (Docker). No code snippets are sent to a third party—ideal for NDAs and code in regulated industries.
- Multi-utilisateurs
- Unlike a purely local extension, Tabby is a server: a single GPU can serve an entire team over the network.
- Telemetry can be disabled
- Tabby reports anonymous statistics by default; one environment variable disables everything. The source code you type is never transmitted.
#Prerequisites
This guide gets you to the model. The kit gets you to the coding copilot in your editor.
- Lifetime online access
- PDF + files
- Lifetime updates
- Docker
- Tabby runs in a container. Docker Desktop on Windows/macOS, Docker Engine on Linux.
- GPU NVIDIA (recommended)
- With the NVIDIA Container Toolkit for CUDA access. Tabby also runs on the CPU, but inline latency becomes uncomfortable.
- ~6 GB of VRAM
- Enough for a 1B–3B completion model in quantized form. More VRAM means a larger model or more developers in parallel.
- An editor
- VS Code or a JetBrains IDE (IntelliJ, PyCharm, WebStorm, etc.). The Tabby extension exists for both.
#Which completion model
Inline autocompletion has one constraint that chat does not: latency. The suggestion must appear within a few hundred milliseconds, or you will type faster than it can respond. So compact Fill-In-the-Middle (FIM) models are preferred, in the -base variant, rather than the large instruct models used for chat.
| VRAM | Completion model | Note |
|---|---|---|
| 4–6 GB | StarCoder2 3B | Responsive, multilingual, a good Tabby default. |
| 6–8 GB | Qwen2.5-Coder 1.5B / 3B (base) | Excellent FIM, very fast, FR/Python/TS. |
| 8–12 GB | Qwen2.5-Coder 7B (base) | The definitive FIM 2026 reference (4.7 GB): sharp suggestions, still-low latency. |
| 12 GB+ | StarCoder2 7B / Qwen-Coder 7B | To share one server among several developers. |
#1. Launch Tabby in Docker
We start the Tabby server with GPU access, a persistent volume for data, and a selected completion model. Port 8080 exposes the web interface and API.
- 01Check startupOn the first launch, Tabby downloads the model (a few minutes). Follow the logs with docker logs -f tabby until the message indicates that the server is listening on port 8080.
- 02Open the web interfaceGo to http://localhost:8080. The first visit prompts you to create an administrator account (email + password). It's local: these credentials remain in your ~/.tabby volume.
- 03Test completionThe interface's “Playground” tab lets you verify that the model responds before connecting an editor.
#2. Create the access token
Editors authenticate with the Tabby server using a token. This lets a team see who is connecting and revoke access without breaking everything.
- 01Go to settingsIn the web interface (http://localhost:8080), open the accounts/security section. The current user’s token is displayed there.
- 02Copy tokenRetrieve the string (often prefixed with auth_). This is what the VS Code and JetBrains extensions will request.
- 03Regenerate if neededIf a token is leaked or a collaborator leaves, regenerate the token from the interface: the old one is immediately invalidated.
#3. Connect VS Code
- 01Install the extensionExtensions (Ctrl+Shift+X) → search for “Tabby” (TabbyML editor) → Install.
- 02Enter the endpointOn the first launch, the extension asks for the server URL. Enter http://localhost:8080 (or the shared server's IP address on the network).
- 03Paste the tokenEnter the access token copied in the previous step. The Tabby icon in the status bar should turn green (connected).
- 04CoderType code: the grayed-out suggestion appears after the cursor. Press Tab to accept it; keep typing to ignore it.
#4. Connect JetBrains
The same configuration works across the entire JetBrains lineup: IntelliJ IDEA, PyCharm, WebStorm, GoLand, Rider, etc. The Tabby extension is a single plugin compatible with the platform.
- 01Install the pluginSettings → Plugins → Marketplace → search for “Tabby” → Install, then restart the IDE.
- 02Configure the connectionSettings → Tools → Tabby: enter the endpoint (http://localhost:8080) and access token.
- 03CheckThe Tabby indicator in the status bar confirms the connection. Inline suggestions appear as they do in VS Code.
#5. Point Tabby to your Ollama
If you already run Ollama for Cline or Aider, there is no need to load a second inference engine: Tabby can use a Ollama backend through its OpenAI-compatible protocol. This lets you share a single model server for completion and chat.
On the Tabby side, declare the HTTP backend instead of a local model. Configure it in the data volume's config.toml file (~/.tabby/config.toml), where you point the completion engine to the Ollama API.
#6. Share the GPU among multiple developers
That's Tabby's decisive advantage over a purely local extension: one server, one GPU, multiple developers. Install Tabby on a machine equipped with a GPU (a dedicated workstation or team server), and each developer points their editor to that network address.
- Expose on the network
- Port 8080 must be accessible from the developers' workstations. Internally, the server's local IP is sufficient (http://192.168.x.x:8080).
- One token per dev
- Create one account/token per developer in the interface. You retain traceability and can revoke access individually.
- GPU sizing
- A compact completion model (1.5B–3B) serves several developers in parallel without becoming saturated. Since completions are brief, requests rarely overlap.
- Reverse proxy + HTTPS
- For access outside the LAN, put Tabby behind a reverse proxy (Caddy, Nginx) with TLS. Never expose port 8080 bare to the Internet.
#Performance tuning and troubleshooting
- Too much latency
- Drop down one model tier (3B → 1.5B), verify that --device cuda is active (not a CPU fallback), and limit the maximum suggestion length.
- No suggestion
- Check the extension’s status icon (correct token/endpoint) and that the container is running (docker ps). The web Playground isolates whether the problem is on the server side or in the editor.
- Off-topic suggestions
- Almost always an incorrect FIM prompt_template or an instruct variant instead of -base. Switch back to a base model and the correct FIM template.
- Repository indexing (RAG)
- Tabby can index your code for more context-aware suggestions. Enable it in the interface if the GPU has headroom; disable it if latency suffers.
- GPU memory saturated
- If several models coexist (Tabby + Cline via Ollama), quantify the KV cache and monitor VRAM with nvidia-smi.
#Frequently asked questions
Does Tabby replace Cline or Copilot?+
Do you need a GPU for Tabby?+
Is my code sent anywhere?+
Can I reuse my existing Ollama?+
How many developers can a single Tabby server serve?+
#Conclusion
You now have a 100% self-hosted code-completion server: Tabby in Docker on your GPU, a StarCoder or Qwen-Coder model in Fill-In-the-Middle mode, an access token, and Copilot-style grayed-out suggestions in VS Code and JetBrains—with the added benefit that you can serve an entire team from a single GPU. Combined with Cline for chat and agent workflows, the picture is complete and your code never leaves your infrastructure. If you want to avoid hours of configuration (FIM templates, Ollama configs, role-specific models), the paid “Local Code Copilot” guide provides a turnkey Ollama + Cline + Aider package, ready to complement Tabby for inline completion.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.