Path 02 · For coding

The developer — « Local Copilot, 100% offline, 0 cloud latency. »

Qwen2.5 Coder 32B connected to VSCode through Continue.dev. Autocomplete, chat, refactoring across an entire file — without a single line of your code leaking to Microsoft, OpenAI, or Anthropic. Does your PR contain an API key? It stays there.

Total duration
~30 min
Required level
Comfortable with a terminal
Cost
0 €
Machine type
GPU with 16 GB VRAM or 32 GB+ Mac M-series
↑
Our promise

At the end of this process, you'll have a llama.cpp server running in the background with Qwen2.5 Coder 32B in Q4 quantization. Continue.dev will be connected to VSCode with autocomplete (FIM) and sidebar chat. Latency will be under 100 ms, and autocomplete quality will be comparable to Copilot, sometimes better on recent code.

Who it's for this path is for you

We'd rather tell you early than let you waste 30 minutes for nothing.

✓ This is for you if
  • ✓You code every day and are tired of copying your snippets into a web UI.
  • ✓Your employer bans GitHub Copilot, ChatGPT, and Cursor — for good reasons.
  • ✓You're working with proprietary code, under an NDA, or with API secrets in plain text.
  • ✓You have a RTX 4070 Ti+, RTX 4090, RTX 5090, or a MacBook Pro M2/M3/M4 with 32 GB+.
  • ✓The 200 ms ping to cloud APIs breaks your rhythm.
✕ Look elsewhere if
  • ·You have 8 GB of VRAM: it's workable, but with a 7B model, quality is well below Copilot.
  • ·You code 1h a week: a free account with a cloud provider will be less of a hassle.
  • ·You want general assistant chat: see the Curious path, which is simpler.

The path in 4 steps

Each step points to a detailed guide you can read alongside it. The order is optimized: don’t skip a step the first time.

Total · ~30 min
  1. 1
    Step 01·12 min·Advanced

    Compile llama.cpp with GPU acceleration

    Build from source for maximum tokens/sec. On NVIDIA: CUDA. On Apple Silicon: Metal. On AMD: ROCm. It’s 15 to 30% faster than Ollama.

    Ready-to-use llama-server binary capable of serving an OpenAI-compatible API.
    Read the guide →
  2. 2
    Step 02·6 min·Intermediate

    Download Qwen2.5 Coder 32B (Q4)

    The best open-weight coding model in its class. Q4_K_M fits in ~19 GB of VRAM while retaining 95% of FP16 quality. Download from Hugging Face.

    The model is served on localhost:8080 in full chat completions format.
    Read the guide →
  3. 3
    Step 03·8 min·Intermediate

    Connect Continue.dev to VSCode

    Install the Continue extension from the marketplace, edit ~/.continue/config.json to point to your local endpoint, and configure the FIM model for autocompletion.

    Tab to accept a suggestion, Cmd+L to open chat for the selection.
    Read the guide →
  4. +
    Step bonus·4 min·Advanced

    Bonus — Helping from the command line

    For larger tasks (refactoring a module, generating tests), Aider lets you edit files in natural language from the terminal, with automatic Git commits.

    An agentic CLI connected to the same backend, complementary to Continue.
    Read the guide →

Recommended models

Three choices tested for this path—from lightest to most capable. Click for the detailed specifications (VRAM, context, benchmarks).

Referenced model
qwen2.5-14b

The versatile one. 14B = 9 GB in Q4, excellent code quality, fits on 12 GB of VRAM.

The all-rounder
Mistral Small 3
Mistral AI · 24B · 14 GB in Q4

A good compromise if you have 16 GB VRAM, with broader capabilities.

View details →
Referenced model
llama-3.3-70b

The best choice if you have a 64 GB Mac or a 48 GB GPU. Higher latency but top-tier quality.

Frequently asked questions

The real questions we get by email and on Mastodon. If yours is missing, open a ticket.

For line-by-line autocomplete, Qwen2.5 Coder 32B is very close, and sometimes better on recent code (post-2024). In chat, Copilot Chat (with GPT-4o behind it) remains a notch ahead. But you retain control of your data.
Once this process is complete
Next path

The confidential pro — « My files never leave my workstation. »

Juriste, médecin, RH, consultant, expert-comptable, notaire — vos documents sont sous secret professionnel ou NDA. Un LLM local couplé à un RAG vous laisse les interroger, résumer,…

→
The Local Copilot Kit

Replace GitHub Copilot and Cursor with a code assistant that runs 100% on your machine — the reference guide, configs included.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

A question, a typo, a bug?

This guide evolves with every model release. Your feedback is the raw material.

Prices in euros (€) are French market prices including VAT, checked by QuelLLM. US prices differ: the Amazon buttons show the current US price.