Intermediate 13 minCode

Best local LLM for coding: Devstral, Qwen3-Coder

The best local LLM for coding in 2026 is no longer a rhetorical question: Devstral, Qwen3-Coder, and the new Qwen 3.5 / 3.8 generation now seriously rival cloud assistants, including on agentic tasks. This guide compares the available open-weight coding models, their VRAM requirements, their real-world generation quality, and provides a clear recommendation based on your GPU. No abstract ranking: practical guidance based on the machine you have.

By Mohamed Meguedmi·Update 2026-08-27·Tested on Windows, macOS, and Linux

#Why use a local LLM for coding in 2026

Coding with an AI assistant has become second nature. But sending proprietary code, secrets, internal paths, or simply intellectual property to a cloud service remains a problem—for freelancers under an NDA, companies subject to the GDPR, and solo developers who simply want to stay in control.

The good news: since 2025, open-weight coding models have made up much of the gap. Devstral reaches SWE-bench scores worthy of the best cloud models, Qwen3-Coder holds its own against Claude for refactoring, and even a Qwen 3.5 9B on a RTX 3060 does useful work every day. The hardware entry cost has dropped, while quality has improved.

i
What we mean by “best local LLM for coding in 2026”
We're talking about open-weight models (Apache 2.0 or MIT), runnable on consumer GPUs or Mac Apple Silicon, without sending a single token to a third-party server. No “open source but API-only,” and no noncommercial license.

#The right selection criteria

The Local Copilot Kit

Devstral or Qwen3-Coder: your choice is made. The Local Copilot Kit puts it to work in thirty minutes (ch. 2), checks that it is the right fit for your VRAM (ch. 5), and has you measure your machine instead of trusting the spec sheet (ch. 18).

  • Lifetime online access
  • PDF + files
  • Lifetime updates

A leaderboard benchmark does not tell the whole story. For real-world use, here is what matters.

Generated code quality
The ability to write a patch that compiles and passes the tests, not just code that “looks right.” HumanEval/MBPP give an indication; SWE-bench Verified is more representative of real-world tasks.
Fill-in-the-Middle (FIM) support
Essential for IDE autocomplete. Not all models support it—Devstral is weak at it, and for pure FIM, qwen2.5-coder:7b-base remains the 2026 reference.
Context window
To understand a 2000-line file or an entire project, allow for at least 32k tokens. Qwen3-Coder goes up to 256k, Devstral to 128k.
Inference speed
A dense 30B model outputs 15-25 tokens/s on a RTX 4090. An MoE such as Qwen3-Coder 30B-A3B outputs 60-80 tokens/s at equivalent quality—the difference changes everything at the keyboard.
License
Apache 2.0 and MIT leave you free to use them commercially. Be cautious with Codestral (non-commercial) or older Code Llama models (restrictive Meta license).
Supported languages
Most good models handle Python, JS/TS, Go, Rust, Java, and C/C++ correctly. For PHP, Ruby, or Swift, check the benchmarks by language.

#Devstral, the agentic specialist (Mistral AI)

Devstral is Mistral’s coding model, specifically trained for agentic work—that is, work with an agent such as Aider, OpenHands, or SWE-agent that iterates on code, runs tests, reads the output, and fixes issues. On SWE-bench Verified, Devstral Small ranks among the best open-weight models, at a level that would have been unattainable a year ago.

Recommended variant
Devstral Small (~24B, dense). Apache 2.0. Available on Hugging Face and Ollama.
VRAM in Q4_K_M
About 14 GB. Easily fits on a RTX 4080 16 GB or a 24 GB M-series Mac. Possible but tight on 12 GB with reduced context.
Context window
128k tokens. More than enough to ingest a medium-sized repo.
Strength
Excellent for multi-step agentic workflows: read a bug report, navigate the code, write a patch, and pass the tests.
Weakness
Not designed for FIM (line-by-line autocompletion). If you want to use it in Continue.dev for completion, switch to qwen2.5-coder:7b-base for that part.
Install Devstral with Ollama
ollama pull devstral:24b
ollama run devstral:24b "Écris un parseur CSV en Rust qui gère les guillemets échappés."

#Qwen3-Coder, the Swiss Army knife (Alibaba)

Qwen3-Coder is Alibaba’s 2025 generation of the Coder family, built on a Mixture-of-Experts (MoE) architecture. It’s the model many consider the state of the art in open-weight code generation in 2026, across all formats. Available under the Apache 2.0 license.

Consumer variant
Qwen3-Coder 30B-A3B: 30 billion total parameters, but only 3 billion activated per token. In practice, that delivers the quality of a dense 14–22B model with the speed of a 3B.
VRAM in Q4_K_M
About 18 GB for the 30B-A3B. Fits exactly on a RTX 4090 24 GB, or on a 24-32 GB M-Pro/M-Max Mac.
Frontier variant
Qwen3-Coder 480B-A35B. Claude/GPT-5 level for coding, but limited to Mac Studio Ultra systems with 256–512 GB or multi-GPU H100 setups.
Context
256k tokens natively, extensible. You can literally paste an entire repo into it.
Strengths
Excellent at multi-file refactoring, handles “niche” languages well (Elixir, Zig, OCaml), very good for agentic workflows, and the MoE makes it exceptionally fast.
Weakness
The 30B-A3B remains an MoE: if your VRAM is tight, the memory overhead versus a dense 9B is noticeable. For 12 GB, stick with a Qwen 3.5 9B (Q8 if necessary).
Install Qwen3-Coder 30B-A3B
ollama pull qwen3-coder:30b
ollama run qwen3-coder:30b "Refactor cette fonction pour utiliser async/await sans changer le comportement."
→
Why the A3B MoE is so fast
With 3B active parameters per token, the computation required for each generation is that of a 3B—but with the “knowledge” of a 30B. On a RTX 4090, expect 60 to 80 tokens/second, whereas a dense 32B tops out at 25-30. In daily use, that's the difference between “I type with the model” and “I wait for the model.”

#Qwen 3.5, the safe choice for small configurations

The Qwen 3.5 family (2B, 4B, 9B) is the most versatile option for modest setups in 2026. Apache 2.0, a large context (up to 256k), multimodal support at the mid-range sizes, and variants designed for 4 to 12 GB of VRAM. For inline autocompletion (FIM) alone, qwen2.5-coder:7b-base remains the reference for this specific use case.

Qwen 3.5 4B
Q4 VRAM ≈ 3.4 GB (ollama run qwen3.5:4b). Ideal for RTX 3060 8 GB, RTX 4060, and 16 GB M-series MacBook Air. The new small default model, decent for coding chat.
Qwen 3.5 9B
Q4 VRAM ≈ 6.6 GB (ollama run qwen3.5:9b). THE 8 GB choice of 2026: 256k context, vision, and clearly better quality than the 4B. The daily driver for 8-12 GB cards.
Qwen 3.5 9B in Q8
VRAM ≈ 11 GB (ollama run qwen3.5:9b-q8_0). The maximum-quality slice, right at 12 GB of VRAM (RTX 3060 12 GB, RTX 4070).
Autocompletion (FIM)
Qwen2.5-Coder 7B base remains THE FIM reference in 2026: ollama run qwen2.5-coder:7b-base (≈ 4.7 GB). Keep it for VS Code's tabAutocomplete (Continue.dev, Tabby).
i
Qwen 3.5 vs Qwen3-Coder
Qwen 3.5 (dense, general-purpose) is recommended for smaller setups (≤16 GB VRAM), with qwen2.5-coder:7b-base as backup for FIM. Qwen3-Coder (code-specialist MoE) targets more generously equipped machines (24 GB+) and agentic work. Both can coexist.

#Notable alternatives

gpt-oss 20B (OpenAI)
Open-weight OpenAI model, quantized with MXFP4, very fast, 131k context. VRAM ≈ 14 GB (ollama run gpt-oss:20b). A good compromise on 16 GB if you want a responsive, code-oriented generalist.
GLM 4.7 Flash (Zhipu AI)
MoE 30B-A3B, MIT-licensed, ≈ 19 GB in Q4 (ollama run glm-4.7-flash). Very solid for technical chat, debugging, and agents. Relevant if you mix French and code—the explanations naturally come out in French.
Mistral Small 24B
Dense general-purpose model, ≈ 14 GB in Q4 (ollama run mistral-small). Good in French, useful as a secondary model for documentation and explanations on a 16 GB setup.
Codestral 22B (Mistral)
Technically correct, but Mistral is a non-production license: prohibited in a work environment. Exclude it except for strictly personal use or research.
DeepSeek-Coder V2, Code Llama, StarCoder 2
These 2023–2024 baselines are outperformed in quality by everything above. Skip them: a Qwen 3.5 9B or a Granite 4.2 8B does better with less VRAM.

#Which one to choose based on your GPU

Pragmatic recommendation based on available VRAM. The models mentioned use Q4_K_M quantization, the best quality-to-memory tradeoff for code.

8 GB (RTX 3060 8 GB, 4060, 5050, 5060)
Qwen 3.5 9B (256k ctx, vision). Already highly usable for autocompletion and technical chat. For pure FIM, add qwen2.5-coder:7b-base. 16-32k context.
12 GB (RTX 3060 12 GB, 4070, 5070)
Qwen 3.5 9B in Q8 (11 GB) as a daily driver, or Gemma 4 12B (7.6 GB, multimodal). Devstral works in Q4 with reduced context.
16 GB (RTX 4080, 4070 Ti Super, 5070 Ti, 5080)
Devstral 24B for agentic tasks, gpt-oss 20B or Mistral Small 24B as general-purpose models, and qwen2.5-coder:7b-base for FIM. This is the first configuration where you really have a choice.
24 GB (RTX 3090, 4090, RX 7900 XTX)
Qwen3-Coder 30B-A3B (code) or Qwen 3.8 27B (general-purpose, closest to a Copilot) as your daily driver. Devstral as an agentic backup, GLM 4.7 Flash for agents. This is the setup where local rivals the cloud for everyday use.
32 GB (RTX 5090) or an M-series Mac with 36–48 GB
Qwen3-Coder 30B-A3B in Q8 (32 GB), or Qwen 3.6 35B-A3B (23 GB, fast MoE, the safe bet in this tier). Excellent comfort, 128k+ context. Aim for the top with no compromises.
Mac Studio Ultra 128 GB+
Qwen3-Coder 30B-A3B in full-quality Q8 with the complete 256k context, or Qwen3-Coder’s frontier variants (480B-A35B) if you are targeting API-level performance. This is the only accessible way, outside a data center, to run a frontier code model locally.

#Connect all of this to VS Code

The simplest setup: Ollama listens on http://localhost:11434, and Continue.dev (the VS Code extension) natively supports this endpoint. Here is a minimal configuration combining qwen2.5-coder:7b-base for autocompletion (fast, FIM) and Qwen3-Coder for chat (powerful).

Retrieve the models
ollama pull qwen2.5-coder:7b-base
ollama pull qwen3-coder:30b
~/.continue/config.json (excerpt)
{
  "models": [
    {
      "title": "Qwen3-Coder 30B (chat)",
      "provider": "ollama",
      "model": "qwen3-coder:30b",
      "apiBase": "http://localhost:11434"
    }
  ],
  "tabAutocompleteModel": {
    "title": "Qwen2.5-Coder 7B base (autocomplete)",
    "provider": "ollama",
    "model": "qwen2.5-coder:7b-base",
    "apiBase": "http://localhost:11434"
  }
}

For agentic workflows (multi-file refactoring, bug resolution), Aider in the CLI is particularly effective with Devstral:

Aider + Devstral
pip install aider-chat
export OLLAMA_API_BASE=http://localhost:11434
aider --model ollama/devstral

#Common tips and pitfalls

Context too short by default on Ollama
Ollama defaults to a 2048-token context limit. For code, that is ridiculous. Configure num_ctx through a Modelfile or through Continue.dev settings, and increase it to at least 16k or 32k.
FIM vs chat—don't mix them up
Devstral and Qwen3-Coder are not optimized for FIM. If you use them for autocompletion, you’ll get odd results (chat-style responses in the middle of a function). Always use a FIM-friendly model (qwen2.5-coder:7b-base) for tabAutocomplete.
Overly aggressive quantization
For code, Q3 and lower visibly degrade quality (broken syntax, invented identifiers). Stay at Q4_K_M minimum. Q5_K_M if you have the VRAM.
Performance collapses after a few minutes
Ollama unloads inactive models after 5 min. If you code in bursts, start Ollama with OLLAMA_KEEP_ALIVE=1h to keep the model warm.
Model that always responds in English
Add a system prompt in Continue.dev or in the Modelfile: “Always respond in French; code and comments in English.” Qwen 3.5 / 3.8 and Devstral follow this instruction without a fuss.
Qwen 3.8 27B that overthinks
With its default reasoning setting, Qwen 3.8 27B tends to overthink simple tasks and increase latency. For everyday coding, set its reasoning effort to low (or turn off thinking): you gain responsiveness with no noticeable loss on common tasks.

#Go further

You've chosen your coding model and it's running. Here are a few ways to take it further:

Local copilot with Continue.dev
The detailed guide to configuring Continue.dev in VS Code with Ollama, separate autocomplete and chat models, and keyboard shortcuts.
Use Ollama in Claude Code and Cursor
To connect Devstral or Qwen3-Coder to Cursor or Claude Code through Ollama's OpenAI-compatible endpoint, and use these high-end IDEs without the cloud.
Choose your quantization (Q4, Q5, Q8, FP16)
To understand exactly what you lose or gain when moving from Q4_K_M to Q5_K_M on a code model, and when increasing quantization makes sense.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.