Twinny: 100% code autocomplete locale
Twinny is a free, open-source VS Code extension (MIT license) that runs code completion on your own machine via Ollama, LM Studio, or llama.cpp. For the current line, a small base model trained for fill-in-the-middle, such as qwen2.5-coder:1.5b-base, responds in under a second. No subscription is required for individuals; billing applies only to teams of more than five developers.
Twinny is an editor extension that provides local code completion: it suggests what comes next based on a model served by your own machine. This is the most latency-sensitive category in the entire ecosystem — a suggestion that arrives in two seconds is useless — which is why model selection follows rules opposite to those used elsewhere. This guide covers the settings that make a difference and what the extension has gained since its early days.
#What Twinny does
Two features have historically coexisted in the extension: inline completion, which displays a grayed-out suggestion while you type, and a sidebar chat for asking questions about selected code. Both rely on a local model served by Ollama, LM Studio, or llama.cpp, and nothing leaves the machine by default.
The value is obvious in a context where code must not leave the premises: a firm, a government agency, or client code covered by a confidentiality agreement. The argument is not cost—it is that the question “where does my code go?” has a verifiable answer. The project describes itself as the AI assistant for VS Code that stays inside your network.
#The detail that changes everything: mid-sequence fill-in
This guide gets you to the model. The kit gets you to the coding copilot in your editor.
- Lifetime online access
- PDF + files
- Lifetime updates
Completing code isn’t the same as continuing text. When you write in the middle of a function, the model must account for what comes before AND what comes after. This capability has a name—fill-in-the-middle, or FIM—and not all models support it: the official documentation is explicit on this point, the completion provider must use a model trained for FIM, and each model family uses distinct tokens to mark the prefix, suffix, and completion area.
This is the number-one source of disappointment. An excellent code-focused chat model without this capability will produce suggestions that ignore the rest of the file and close braces that are already there. Another point many people overlook: the documentation recommends a base model rather than an instruct variant, because a base model continues raw text more reliably than it handles an empty suffix, especially when there is nothing after the cursor.
#Settings that make the experience smooth
Four settings, documented by the project, account for most of the perceived latency improvement. They live in the extension's settings (twinny. prefix) and can all be adjusted without restarting Ollama.
| Parameter | Default value | Effect |
|---|---|---|
| twinny.debounceWait | 300 ms | Delay after the last keystroke before sending the request; increasing it reduces the number of calls |
| twinny.numPredictFim | 512 tokens | Maximum number of tokens generated per suggestion; reducing it speeds up the response |
| twinny.contextLength | 100 lines | Amount of code sent around the cursor; reducing it lowers latency |
| twinny.fileContextEnabled | disabled by default | Send excerpts from other related open files (up to 3); useful on a large project but adds latency |
- 01Limit the number of generated tokensA useful completion is one or two lines. The documentation confirms it: if suggestions arrive late, the first thing to lower is numPredictFim, not the model.
- 02Adjust the trigger delayTriggering on every keystroke overloads the server; the default 300 ms delay (debounceWait) is sufficient in most cases and cuts the number of calls in half.
- 03Keep the model loaded on the Ollama sideIf the server unloads the model after a few minutes of inactivity, the first suggestion after a pause arrives too late. Increase the keep-alive duration on the Ollama side (keep_alive).
- 04Reduce the context window sentThe documentation says so explicitly: if suggestions seem delayed, lower contextLength (100 lines by default before and after the cursor) or leave fileContextEnabled disabled, which is its default setting. This parameter only adds excerpts from other open files when you enable it yourself.
#What the model actually sees
A suggestion doesn't rely only on the prefix and suffix around the cursor. The documentation details the sources Twinny assembles before building the prompt: surrounding lines (contextLength), the latest changes made in the file as small diffs (recentEditsEnabled, enabled by default, up to 6 changes), the symbols and signature of the current call provided by the language server (lspContextEnabled, enabled by default), and optionally excerpts from other related open files (fileContextEnabled, disabled by default).
- Recent changes
- A half-finished rename or a pattern applied manually to a file propagates into subsequent suggestions, as if it understood what you are doing.
- IntelliSense context
- The model receives the names and signature of the call containing the cursor, supplied by the editor's language server.
- Related files
- Disabled by default: enable it only for a large project, because it adds latency for variable gains.
Another useful practical detail: first write a one-line comment describing what comes next, then pause; this gives multi-line completion a clear objective—the advice provided by the documentation itself for getting better suggestions. Clearly naming your variables and functions remains, before any tuning, the factor with the greatest impact on relevance.
#Which models, and why so small
| Usage | Recommended size | Why |
|---|---|---|
| Inline completion (FIM) | 1.5 to 3 billion, base model | Latency comes first: qwen2.5-coder:1.5b-base responds in under a second on most machines, according to the official documentation |
| Completion on a high-end card | 7 billion | Feasible if the card is fast and the context is reduced (low contextLength) |
| Discussion sidebar | 7 to 14 billion, instruct model | Here, we’re willing to wait for a well-constructed answer; a base model isn’t suited to conversation |
This asymmetry is always surprising: we’re used to bigger being better. For completion, a 1.5-billion-parameter model trained for fill-in-the-middle and served in base form beats a general-purpose 14-billion-parameter model in practice because it responds before you’ve finished thinking. The extension lets you configure a model by function—one for FIM, another for discussion—and that is the right setup: the two have neither the same size nor the same training.
#Beyond completion: what Twinny does in 2026
The extension has grown considerably since its beginnings as a simple completion plugin. Its official page now lists inline editing (describe a change and review it as a diff in the editor), a workstation index combining keyword and vector search, code review for the working tree, a branch, or a GitHub pull request, commit message generation from staged changes, and terminal command generation with automatic error correction.
All these features remain tied to remappable VS Code commands and customizable prompt templates. On the provider side, the list has expanded to hosted APIs (OpenAI, Anthropic, Mistral, DeepSeek, OpenRouter, Gemini, Groq, Cohere, Perplexity) in addition to local engines (Ollama, LM Studio, llama.cpp, Oobabooga, LiteLLM, Open WebUI)—but you are not required to use any of them: 100% local usage remains the default configuration.
#Free for individuals, paid for teams
For individual use, Twinny remains what it has always been: an MIT-licensed extension with no telemetry or mandatory connection. Billing only appears beyond five developers sharing a team gateway (twinny-server), which centralizes provider access for the entire organization: free for up to five accounts, then charged per seat per month, with a 30-day trial requiring no bank card. For a solo developer with Ollama locally, none of these conditions apply, and the extension remains fully usable without ever creating an account.
#Twinny or an editing agent
| Need | Tool |
|---|---|
| Finish the current line without thinking about it | Twinny — inline completion (FIM) |
| Modify multiple files based on instructions | An editing agent integrated into the editor (Roo Code, Cline) |
| Run a complete task without supervision | An autonomous agent in a terminal or container |
| Reviewing code before a commit | Twinny’s code review feature, or a dedicated conversational assistant |
The two use cases are not mutually exclusive. Many developers keep Twinny running constantly for inline completion—invisible, fast, and powered by a small dedicated model—and switch to an editing agent such as Roo Code or Cline for tasks that affect multiple files, using a larger model they are willing to wait a few seconds for.
- Roo Code: the local coding agent in your editor
- Cline + Ollama: 100% local coding agent in VS Code
- Review and revise code with a local LLM
- Best local LLM for coding: comparison
- Source: Twinny’s official GitHub repository
- Source: the FIM completion documentation
- Source: Visual Studio Marketplace listing
#FAQ
Is Twinny free?+
Which model for inline completion?+
Why do my suggestions ignore the rest of the file?+
How can you reduce suggestion latency?+
Do you need a high-end graphics card?+
Twinny or an editing agent such as Roo Code or Cline?+
Does Twinny do anything besides code completion?+
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.