Home › Catalog › Best local LLM for coding in 2026

Best local LLM for coding in 2026

Ranking updated on 09/10/2026

Ranking of open-weight LLMs specialized in or capable of code generation, refactoring, and autocompletion. Preference is given to models evaluated on HumanEval/MBPP, with a context of ≥ 32k tokens to cover entire files, and a permissive license.

⚡ The complete setup that worksCline + Aider + tuned Modelfiles, ready in 30 minutes — the Local Copilot Kit →

Ranking

1

🇺🇸 Laguna XS.2

Poolside · 33B parameters · Apache 2.0 · 131,072 tokens ctx

MoE 33B/3B active parameters, Apache 2.0, specializing in agentic coding. 68.2% SWE-Bench Verified, 128k ctx. Runs on a 36 GB Mac. Released April 28, 2026.

Why this ranking Code-specialized model, 131,072-token context. Permissive commercial license.
ollama run laguna-xs.2
Q4 VRAM
19 GB
35 GB in Q8
2

🇫🇷 Devstral Small 2 24B

Mistral AI · 24B parameters · Apache 2.0 · 256,000-token context

24B coding specialist, Apache 2.0. 72.2% SWE-Bench. 256k ctx, FR lab.

Why this ranking Code-specialized model, 256,000-token context. Commercially permissive license.
ollama run devstral-small2:24b
Q4 VRAM
14 GB
26 GB in Q8
3

🇺🇸 Laguna XS 2.1

Poolside · 33B parameters · OpenMDW 1.1 · 256,000-token context

MoE 33B / 3B active (Poolside), multilingual agentic coding. 256k context, open OpenMDW license. Runs on a Mac with 43 GB. Release: June 2026.

Why this ranking Code-specialized model with a 256,000-token context. License must be verified for production use.
ollama run laguna-xs-2.1
Q4 VRAM
19 GB
35 GB in Q8
4

🇨🇳 Qwen3-Coder 30B-A3B

Alibaba · 30B parameters · Apache 2.0 · 262,144-token context

MoE 30B (3.3B active parameters) specialized in agentic coding. Very fast locally, native 256k ctx, the benchmark for 16–24 GB via Ollama.

Why this ranking Code-specialized model, 262,144-token context. Commercially permissive license.
ollama run qwen3-coder:30b
Q4 VRAM
19 GB
35 GB in Q8
5

🇨🇳 Qwen 2.5 Coder 32B

Alibaba · 32B parameters · Apache 2.0 · 131,072 tokens ctx

The open-weight code reference from late 2024, now surpassed by the Qwen3-Coder / Devstral generation.

Why this ranking Code-specialized model, 131,072-token context. Permissive commercial license.
ollama run qwen2.5-coder:32b
Q4 VRAM
19 GB
35 GB in Q8
6

🇨🇳 Qwen 2.5 Coder 14B Instruct

Alibaba · 14B parameters · Apache 2.0 · 131,072 tokens ctx

Coding 14B. HumanEval 89.6, LiveCodeBench 37.1. VRAM sweet spot for self-hosted coding.

Why this ranking Code-specialized model, 131,072-token context. Permissive commercial license.
ollama run qwen2.5-coder:14b
Q4 VRAM
9 GB
16 GB in Q8
7

🇨🇳 Qwen 3.6 27B

Alibaba · 27B parameters · Apache 2.0 · 262,144-token context

Dense multimodal 27B released April 22, 2026. 262k ctx (1M YaRN). SWE-bench Verified 77.2%.

Why this ranking Code-specialized model, 262,144-token context. Commercially permissive license.
ollama run qwen3.6:27b
Q4 VRAM
16 GB
29 GB in Q8

Comparison table

Rank Model Params Q4 VRAM Context License
#1 Laguna XS.2 33B 19 GB 131 072 Apache 2.0
#2 Devstral Small 2 24B 24B 14 GB 256 000 Apache 2.0
#3 Laguna XS 2.1 33B 19 GB 256 000 OpenMDW 1.1
#4 Qwen3-Coder 30B-A3B 30B 19 GB 262 144 Apache 2.0
#5 Qwen 2.5 Coder 32B 32B 19 GB 131 072 Apache 2.0
#6 Qwen 2.5 Coder 14B Instruct 14B 9 GB 131 072 Apache 2.0
#7 Qwen 3.6 27B 27B 16 GB 262 144 Apache 2.0

Two options depending on the memory you need and your software: Radeon RX 9070 XT 16 GB · GMKtec EVO-X2 64 GB / 1 TB (Ryzen AI Max+ 395) — compare AI hardware →

The Local Copilot kit

You have your model. Now you need to make it work in your editor: the Local Copilot kit connects it to VS Code with Cline (ch. 7), to the terminal with Aider (ch. 8), and to autocomplete as you type (ch. 9)—using the Modelfile tuned for this specific model (ch. 11).

  • Lifetime online access
  • PDF + files
  • Lifetime updates

Free memo

Which coding model should you run on YOUR machine?

Get the memo VRAM → best coding model → Ollama command (one screen, copy and paste). Then switch to the Copilote Local kit for a setup that actually works.

The Local Copilot kit — the Ollama + Cline + Aider configs are ready to paste, with tuned Modelfiles, troubleshooting, and lifetime online access →

No spam. Unsubscribe in 1 click. Your data stays with us (never resold).

Ranking methodology

We filter models whose tags include “code” and that remain ≤ 100B (beyond that, they’re no longer “local”). The ranking combines code specialization (name contains “coder,” “codestral,” “devstral,” or “laguna”), long context, a permissive license, and a 7–32B sweet-spot size.

Criteria considered:

  • HumanEval / MBPP
  • Context ≥ 32k tokens
  • Permissive license
  • IDE support (Cline, Aider)

The scoring is fully transparent: see our methodology for details on VRAM/tokens/sec calculations.

Frequently asked questions

What is the best open-source LLM for coding in 2026?

Our #1 is Laguna XS.2 (33B, Apache 2.0). It combines high HumanEval accuracy, a 131,072-token context, and a permissive license.

How much VRAM do you need for a good coding LLM?

For a 7B Coder model in Q4_K_M, 6–8 GB of VRAM is enough. For a 32B Coder in Q4, plan for 20–24 GB (RTX 4090, 3090, 7900 XTX). Q5 or Q8 climbs quickly — consult the configurator.

Can I use it in commercial production?

Yes, if the license is Apache 2.0 or MIT. Codestral is under Mistral Non-Production License — for personal/research use only. Qwen and DeepSeek are free.

How do you integrate a local LLM into VS Code?

Via the extension Cline, which connects to your local Ollama or LM Studio server. Aider works in the CLI for multi-file refactoring.

Go further

QuelLLM Kits The reference guide by use case
All kits for life — $49