Intermediate 11 minIDE

Twinny: 100% code autocomplete locale

Direct response

Twinny is a free, open-source VS Code extension (MIT license) that runs code completion on your own machine via Ollama, LM Studio, or llama.cpp. For the current line, a small base model trained for fill-in-the-middle, such as qwen2.5-coder:1.5b-base, responds in under a second. No subscription is required for individuals; billing applies only to teams of more than five developers.

Twinny is an editor extension that provides local code completion: it suggests what comes next based on a model served by your own machine. This is the most latency-sensitive category in the entire ecosystem — a suggestion that arrives in two seconds is useless — which is why model selection follows rules opposite to those used elsewhere. This guide covers the settings that make a difference and what the extension has gained since its early days.

By Mohamed Meguedmi·Update 2026-09-29·Tested on Windows, macOS, and Linux

#What Twinny does

Two features have historically coexisted in the extension: inline completion, which displays a grayed-out suggestion while you type, and a sidebar chat for asking questions about selected code. Both rely on a local model served by Ollama, LM Studio, or llama.cpp, and nothing leaves the machine by default.

The value is obvious in a context where code must not leave the premises: a firm, a government agency, or client code covered by a confidentiality agreement. The argument is not cost—it is that the question “where does my code go?” has a verifiable answer. The project describes itself as the AI assistant for VS Code that stays inside your network.

#The detail that changes everything: mid-sequence fill-in

The Local Copilot Kit

This guide gets you to the model. The kit gets you to the coding copilot in your editor.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

Completing code isn’t the same as continuing text. When you write in the middle of a function, the model must account for what comes before AND what comes after. This capability has a name—fill-in-the-middle, or FIM—and not all models support it: the official documentation is explicit on this point, the completion provider must use a model trained for FIM, and each model family uses distinct tokens to mark the prefix, suffix, and completion area.

This is the number-one source of disappointment. An excellent code-focused chat model without this capability will produce suggestions that ignore the rest of the file and close braces that are already there. Another point many people overlook: the documentation recommends a base model rather than an instruct variant, because a base model continues raw text more reliably than it handles an empty suffix, especially when there is nothing after the cursor.

→
The model recommended by the documentation
“Use a small base model on the GPU. qwen2.5-coder:1.5b-base responds in well under one second on most machines” — this is the official recommendation for inline completion.

#Settings that make the experience smooth

Four settings, documented by the project, account for most of the perceived latency improvement. They live in the extension's settings (twinny. prefix) and can all be adjusted without restarting Ollama.

The four settings that control perceived latency
ParameterDefault valueEffect
twinny.debounceWait300 msDelay after the last keystroke before sending the request; increasing it reduces the number of calls
twinny.numPredictFim512 tokensMaximum number of tokens generated per suggestion; reducing it speeds up the response
twinny.contextLength100 linesAmount of code sent around the cursor; reducing it lowers latency
twinny.fileContextEnableddisabled by defaultSend excerpts from other related open files (up to 3); useful on a large project but adds latency
  1. 01
    Limit the number of generated tokens
    A useful completion is one or two lines. The documentation confirms it: if suggestions arrive late, the first thing to lower is numPredictFim, not the model.
  2. 02
    Adjust the trigger delay
    Triggering on every keystroke overloads the server; the default 300 ms delay (debounceWait) is sufficient in most cases and cuts the number of calls in half.
  3. 03
    Keep the model loaded on the Ollama side
    If the server unloads the model after a few minutes of inactivity, the first suggestion after a pause arrives too late. Increase the keep-alive duration on the Ollama side (keep_alive).
  4. 04
    Reduce the context window sent
    The documentation says so explicitly: if suggestions seem delayed, lower contextLength (100 lines by default before and after the cursor) or leave fileContextEnabled disabled, which is its default setting. This parameter only adds excerpts from other open files when you enable it yourself.

#What the model actually sees

A suggestion doesn't rely only on the prefix and suffix around the cursor. The documentation details the sources Twinny assembles before building the prompt: surrounding lines (contextLength), the latest changes made in the file as small diffs (recentEditsEnabled, enabled by default, up to 6 changes), the symbols and signature of the current call provided by the language server (lspContextEnabled, enabled by default), and optionally excerpts from other related open files (fileContextEnabled, disabled by default).

Recent changes
A half-finished rename or a pattern applied manually to a file propagates into subsequent suggestions, as if it understood what you are doing.
IntelliSense context
The model receives the names and signature of the call containing the cursor, supplied by the editor's language server.
Related files
Disabled by default: enable it only for a large project, because it adds latency for variable gains.

Another useful practical detail: first write a one-line comment describing what comes next, then pause; this gives multi-line completion a clear objective—the advice provided by the documentation itself for getting better suggestions. Clearly naming your variables and functions remains, before any tuning, the factor with the greatest impact on relevance.

i
Run the model on another one of your devices
Beyond Ollama, LM Studio, and llama.cpp locally, Twinny offers direct device pairing (“Devices”): your workstation uses the GPU in another of your machines over an encrypted peer-to-peer connection, with no account or relay.

#Which models, and why so small

The rules are reversed compared with the discussion
UsageRecommended sizeWhy
Inline completion (FIM)1.5 to 3 billion, base modelLatency comes first: qwen2.5-coder:1.5b-base responds in under a second on most machines, according to the official documentation
Completion on a high-end card7 billionFeasible if the card is fast and the context is reduced (low contextLength)
Discussion sidebar7 to 14 billion, instruct modelHere, we’re willing to wait for a well-constructed answer; a base model isn’t suited to conversation

This asymmetry is always surprising: we’re used to bigger being better. For completion, a 1.5-billion-parameter model trained for fill-in-the-middle and served in base form beats a general-purpose 14-billion-parameter model in practice because it responds before you’ve finished thinking. The extension lets you configure a model by function—one for FIM, another for discussion—and that is the right setup: the two have neither the same size nor the same training.

i
Measure latency, not quality
To evaluate a code-completion setup, the right metric is the time before the suggestion appears. Beyond roughly half a second, the tool becomes more distracting than helpful, regardless of how relevant its suggestion is.

#Beyond completion: what Twinny does in 2026

The extension has grown considerably since its beginnings as a simple completion plugin. Its official page now lists inline editing (describe a change and review it as a diff in the editor), a workstation index combining keyword and vector search, code review for the working tree, a branch, or a GitHub pull request, commit message generation from staged changes, and terminal command generation with automatic error correction.

All these features remain tied to remappable VS Code commands and customizable prompt templates. On the provider side, the list has expanded to hosted APIs (OpenAI, Anthropic, Mistral, DeepSeek, OpenRouter, Gemini, Groq, Cohere, Perplexity) in addition to local engines (Ollama, LM Studio, llama.cpp, Oobabooga, LiteLLM, Open WebUI)—but you are not required to use any of them: 100% local usage remains the default configuration.

#Free for individuals, paid for teams

For individual use, Twinny remains what it has always been: an MIT-licensed extension with no telemetry or mandatory connection. Billing only appears beyond five developers sharing a team gateway (twinny-server), which centralizes provider access for the entire organization: free for up to five accounts, then charged per seat per month, with a 30-day trial requiring no bank card. For a solo developer with Ollama locally, none of these conditions apply, and the extension remains fully usable without ever creating an account.

#Twinny or an editing agent

Three categories of coding assistance
NeedTool
Finish the current line without thinking about itTwinny — inline completion (FIM)
Modify multiple files based on instructionsAn editing agent integrated into the editor (Roo Code, Cline)
Run a complete task without supervisionAn autonomous agent in a terminal or container
Reviewing code before a commitTwinny’s code review feature, or a dedicated conversational assistant

The two use cases are not mutually exclusive. Many developers keep Twinny running constantly for inline completion—invisible, fast, and powered by a small dedicated model—and switch to an editing agent such as Roo Code or Cline for tasks that affect multiple files, using a larger model they are willing to wait a few seconds for.

#FAQ

Is Twinny free?+
Yes for individual use: the extension is MIT-licensed, open source, with no telemetry or required account, and the models it uses run on your computer. Billing applies only to the shared team gateway used by more than five developers, starting at $6 per seat per month after a 30-day free trial.
Which model for inline completion?+
A small base model trained for fill-in-the-middle (FIM), not a chat model. The official documentation recommends qwen2.5-coder:1.5b-base, which responds in under one second on most machines and remains relevant for the next few lines. An instruct model or a larger one often responds too late to be useful while typing, even if it performs better on a long question.
Why do my suggestions ignore the rest of the file?+
The selected model probably does not support fill-in-the-middle, or it is an instruction-tuned model rather than a base model: the documentation specifies that a base model handles an empty suffix better than an instruction-tuned model. Check the model page and the FIM template configured in the extension.
How can you reduce suggestion latency?+
First lower twinny.numPredictFim (512 tokens by default), then twinny.contextLength (100 lines before and after the cursor by default); leave twinny.fileContextEnabled disabled if the project is large, and keep the model loaded in memory on Ollama between suggestions. These are the settings the official documentation prioritizes for user-perceived latency.
Do you need a high-end graphics card?+
No, and that is the good news: a 1.5-billion-parameter completion model fits in a few gigabytes of VRAM and responds quickly even on modest hardware, provided you stick with a base model dedicated to FIM rather than a general-purpose model.
Twinny or an editing agent such as Roo Code or Cline?+
They are not the same tools. Twinny completes the current line with a small, fast model; an editing agent modifies multiple files based on instructions using a larger model. Many developers use both, with different models for each task.
Does Twinny do anything besides code completion?+
Yes: since its latest versions, the extension has added inline editing by diff, a workstation index (keyword and vector search), code review on the working tree or a pull request, and generation of commit messages and terminal commands.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.