Intermediate 25 minIDE

Tabby: install local code autocompletion self-hosted

Cline and Aider provide chat and agent mode, but they don’t provide the inline “gray text” autocompletion that appears as you type—the hallmark of GitHub Copilot. Tabby fills that exact gap: it’s a self-hosted, open-source code-completion server (Apache 2.0) that runs in Docker on your GPU and provides Fill-In-the-Middle suggestions directly in the editor. This guide installs it step by step, configures a StarCoder or Qwen-Coder model, generates the token, connects VS Code and JetBrains, and shows how a single GPU can serve an entire team—your code never leaves your infrastructure.

By Mohamed Meguedmi·Update 2026-08-27·Tested on Windows 11

#Why Tabby instead of Cline

Tabby and Cline are not competitors: they complement each other. Cline (and Aider) excels at conversational chat and multi-file agent mode, but neither provides line-by-line autocomplete while you type. Tabby does exactly that—and does it very well. The winning combination in 2026: Cline for discussion and refactoring, Tabby for the gray suggestion that appears as you type.

Inline grayed-out text
The suggestion appears directly after the cursor, grayed out. Press Tab to accept, Esc to ignore. Exactly Copilot's UX.
100% self-hosted
The server runs locally (Docker). No code snippets are sent to a third party—ideal for NDAs and code in regulated industries.
Multi-utilisateurs
Unlike a purely local extension, Tabby is a server: a single GPU can serve an entire team over the network.
Telemetry can be disabled
Tabby reports anonymous statistics by default; one environment variable disables everything. The source code you type is never transmitted.
i
Tabby ≠ Cline
If you are looking for chat or an agent that modifies multiple files, Tabby is not it: see the Cline guides instead. Tabby does ONLY inline autocompletion. The two coexist perfectly in the same editor, connected to the same GPU.

#Prerequisites

The Local Copilot Kit

This guide gets you to the model. The kit gets you to the coding copilot in your editor.

  • Lifetime online access
  • PDF + files
  • Lifetime updates
Docker
Tabby runs in a container. Docker Desktop on Windows/macOS, Docker Engine on Linux.
GPU NVIDIA (recommended)
With the NVIDIA Container Toolkit for CUDA access. Tabby also runs on the CPU, but inline latency becomes uncomfortable.
~6 GB of VRAM
Enough for a 1B–3B completion model in quantized form. More VRAM means a larger model or more developers in parallel.
An editor
VS Code or a JetBrains IDE (IntelliJ, PyCharm, WebStorm, etc.). The Tabby extension exists for both.
→
On Mac Apple Silicon
Docker GPU passthrough is not available on macOS. On an M1/M2/M3 Mac, run the native Tabby binary (Metal) downloaded from the GitHub releases instead of the Docker CUDA image. The editor configuration remains the same.

#Which completion model

Inline autocompletion has one constraint that chat does not: latency. The suggestion must appear within a few hundred milliseconds, or you will type faster than it can respond. So compact Fill-In-the-Middle (FIM) models are preferred, in the -base variant, rather than the large instruct models used for chat.

Recommended completion model by VRAM
VRAMCompletion modelNote
4–6 GBStarCoder2 3BResponsive, multilingual, a good Tabby default.
6–8 GBQwen2.5-Coder 1.5B / 3B (base)Excellent FIM, very fast, FR/Python/TS.
8–12 GBQwen2.5-Coder 7B (base)The definitive FIM 2026 reference (4.7 GB): sharp suggestions, still-low latency.
12 GB+StarCoder2 7B / Qwen-Coder 7BTo share one server among several developers.
!
-base variant, never -instruct
For FIM autocompletion, always use the -base variant of the model, not the instruct/chat version. An instruct variant tends to “chatter” (adding comments and explanations) instead of completing the code cleanly. Keep instruct models for Cline.

#1. Launch Tabby in Docker

We start the Tabby server with GPU access, a persistent volume for data, and a selected completion model. Port 8080 exposes the web interface and API.

Start Tabby (GPU NVIDIA)
docker run -d \
  --name tabby \
  --gpus all \
  -p 8080:8080 \
  -v $HOME/.tabby:/data \
  registry.tabbyml.com/tabbyml/tabby \
  serve --model StarCoder2-3B --device cuda
  1. 01
    Check startup
    On the first launch, Tabby downloads the model (a few minutes). Follow the logs with docker logs -f tabby until the message indicates that the server is listening on port 8080.
  2. 02
    Open the web interface
    Go to http://localhost:8080. The first visit prompts you to create an administrator account (email + password). It's local: these credentials remain in your ~/.tabby volume.
  3. 03
    Test completion
    The interface's “Playground” tab lets you verify that the model responds before connecting an editor.
→
Disable telemetry
Add -e TABBY_DISABLE_USAGE_COLLECTION=1 to the docker run command to disable all anonymous statistics reporting. Your code content is never transmitted anyway, but this also disables usage metrics.

#2. Create the access token

Editors authenticate with the Tabby server using a token. This lets a team see who is connecting and revoke access without breaking everything.

  1. 01
    Go to settings
    In the web interface (http://localhost:8080), open the accounts/security section. The current user’s token is displayed there.
  2. 02
    Copy token
    Retrieve the string (often prefixed with auth_). This is what the VS Code and JetBrains extensions will request.
  3. 03
    Regenerate if needed
    If a token is leaked or a collaborator leaves, regenerate the token from the interface: the old one is immediately invalidated.
What the editor needs
Endpoint : http://localhost:8080
Token    : auth_xxxxxxxxxxxxxxxxxxxxxxxx

#3. Connect VS Code

  1. 01
    Install the extension
    Extensions (Ctrl+Shift+X) → search for “Tabby” (TabbyML editor) → Install.
  2. 02
    Enter the endpoint
    On the first launch, the extension asks for the server URL. Enter http://localhost:8080 (or the shared server's IP address on the network).
  3. 03
    Paste the token
    Enter the access token copied in the previous step. The Tabby icon in the status bar should turn green (connected).
  4. 04
    Coder
    Type code: the grayed-out suggestion appears after the cursor. Press Tab to accept it; keep typing to ignore it.
→
Configure triggering
In the extension settings, you can switch between automatic triggering (the suggestion appears on its own) and manual triggering (on demand via a shortcut). Manual mode saves GPU resources if several developers share the server.

#4. Connect JetBrains

The same configuration works across the entire JetBrains lineup: IntelliJ IDEA, PyCharm, WebStorm, GoLand, Rider, etc. The Tabby extension is a single plugin compatible with the platform.

  1. 01
    Install the plugin
    Settings → Plugins → Marketplace → search for “Tabby” → Install, then restart the IDE.
  2. 02
    Configure the connection
    Settings → Tools → Tabby: enter the endpoint (http://localhost:8080) and access token.
  3. 03
    Check
    The Tabby indicator in the status bar confirms the connection. Inline suggestions appear as they do in VS Code.

#5. Point Tabby to your Ollama

If you already run Ollama for Cline or Aider, there is no need to load a second inference engine: Tabby can use a Ollama backend through its OpenAI-compatible protocol. This lets you share a single model server for completion and chat.

FIM model on the Ollama side
# Tirez un modèle de complétion FIM dans Ollama
ollama pull qwen2.5-coder:1.5b-base
# ou
ollama pull starcoder2:3b

On the Tabby side, declare the HTTP backend instead of a local model. Configure it in the data volume's config.toml file (~/.tabby/config.toml), where you point the completion engine to the Ollama API.

~/.tabby/config.toml
[model.completion.http]
kind = "ollama/completion"
model_name = "qwen2.5-coder:1.5b-base"
api_endpoint = "http://localhost:11434"
prompt_template = "<|fim_prefix|>{prefix}<|fim_suffix|>{suffix}<|fim_middle|>"
!
The FIM template depends on the model
Each model family has its own Fill-In-the-Middle tags (StarCoder, Qwen, and DeepSeek differ). An incorrect prompt_template produces absurd suggestions. Check your model’s FIM tags in its model card before locking in the configuration.
i
Localhost as seen from the container
From the Tabby container, http://localhost:11434 does not point to the host’s Ollama. Use http://host.docker.internal:11434 (Docker Desktop) or the host’s IP address on the Docker network; otherwise, Tabby will never reach Ollama.

#6. Share the GPU among multiple developers

That's Tabby's decisive advantage over a purely local extension: one server, one GPU, multiple developers. Install Tabby on a machine equipped with a GPU (a dedicated workstation or team server), and each developer points their editor to that network address.

Expose on the network
Port 8080 must be accessible from the developers' workstations. Internally, the server's local IP is sufficient (http://192.168.x.x:8080).
One token per dev
Create one account/token per developer in the interface. You retain traceability and can revoke access individually.
GPU sizing
A compact completion model (1.5B–3B) serves several developers in parallel without becoming saturated. Since completions are brief, requests rarely overlap.
Reverse proxy + HTTPS
For access outside the LAN, put Tabby behind a reverse proxy (Caddy, Nginx) with TLS. Never expose port 8080 bare to the Internet.
→
The ideal team server
A single mid-range card (12–16 GB) is enough to equip a small team with autocomplete. Compared with a per-seat Copilot subscription, the return on investment of a shared GPU is fast—and all the code stays on your infrastructure.

#Performance tuning and troubleshooting

Too much latency
Drop down one model tier (3B → 1.5B), verify that --device cuda is active (not a CPU fallback), and limit the maximum suggestion length.
No suggestion
Check the extension’s status icon (correct token/endpoint) and that the container is running (docker ps). The web Playground isolates whether the problem is on the server side or in the editor.
Off-topic suggestions
Almost always an incorrect FIM prompt_template or an instruct variant instead of -base. Switch back to a base model and the correct FIM template.
Repository indexing (RAG)
Tabby can index your code for more context-aware suggestions. Enable it in the interface if the GPU has headroom; disable it if latency suffers.
GPU memory saturated
If several models coexist (Tabby + Cline via Ollama), quantify the KV cache and monitor VRAM with nvidia-smi.
Quick diagnostics
# Le conteneur tourne-t-il ?
docker ps | grep tabby
# Logs en direct
docker logs -f tabby
# Le GPU est-il bien vu ?
docker exec tabby nvidia-smi

#Frequently asked questions

FAQ
Does Tabby replace Cline or Copilot?+
Tabby replaces ONLY Copilot's inline “gray text” autocomplete—which is precisely what Cline does not do. Keep Cline for chat and multi-file agent mode. Both coexist in the same editor, connected to the same GPU: Tabby completes as you type, while Cline chats and refactors.
Do you need a GPU for Tabby?+
Strongly recommended. Tabby runs on the CPU, but inline autocompletion requires latency of a few hundred milliseconds: without a GPU, the suggestion arrives after you’ve already typed the line. A NVIDIA card with ~6 GB of VRAM and the NVIDIA Container Toolkit is enough for a compact completion model.
Is my code sent anywhere?+
No. Tabby is self-hosted: the server runs on your machine or your team server, and the code content is never transmitted to a third party. By default, Tabby sends anonymous usage statistics (never code); the TABBY_DISABLE_USAGE_COLLECTION=1 variable disables them completely.
Can I reuse my existing Ollama?+
Yes. Tabby can consume a Ollama backend through its compatible protocol. Declare the HTTP completion engine in config.toml, pointing it to the Ollama API with the correct prompt_template FIM. This lets you share a single model server for completion (Tabby) and chat (Cline).
How many developers can a single Tabby server serve?+
Several, on a single GPU. Since completions are brief and intermittent, a 1.5B–3B model on a 12–16 GB card can handle a small team without saturating it. Create one token per developer for traceability, and expose port 8080 on the LAN (or behind a TLS reverse proxy for remote access).

#Conclusion

You now have a 100% self-hosted code-completion server: Tabby in Docker on your GPU, a StarCoder or Qwen-Coder model in Fill-In-the-Middle mode, an access token, and Copilot-style grayed-out suggestions in VS Code and JetBrains—with the added benefit that you can serve an entire team from a single GPU. Combined with Cline for chat and agent workflows, the picture is complete and your code never leaves your infrastructure. If you want to avoid hours of configuration (FIM templates, Ollama configs, role-specific models), the paid “Local Code Copilot” guide provides a turnkey Ollama + Cline + Aider package, ready to complement Tabby for inline completion.

Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.