🇺🇸 Laguna XS.2
MoE 33B/3B active parameters, Apache 2.0, specializing in agentic coding. 68.2% SWE-Bench Verified, 128k ctx. Runs on a 36 GB Mac. Released April 28, 2026.
ollama run laguna-xs.2
Ranking updated on 09/10/2026
Ranking of open-weight LLMs specialized in or capable of code generation, refactoring, and autocompletion. Preference is given to models evaluated on HumanEval/MBPP, with a context of ≥ 32k tokens to cover entire files, and a permissive license.
⚡ The complete setup that worksCline + Aider + tuned Modelfiles, ready in 30 minutes — the Local Copilot Kit →MoE 33B/3B active parameters, Apache 2.0, specializing in agentic coding. 68.2% SWE-Bench Verified, 128k ctx. Runs on a 36 GB Mac. Released April 28, 2026.
ollama run laguna-xs.2
24B coding specialist, Apache 2.0. 72.2% SWE-Bench. 256k ctx, FR lab.
ollama run devstral-small2:24b
MoE 33B / 3B active (Poolside), multilingual agentic coding. 256k context, open OpenMDW license. Runs on a Mac with 43 GB. Release: June 2026.
ollama run laguna-xs-2.1
MoE 30B (3.3B active parameters) specialized in agentic coding. Very fast locally, native 256k ctx, the benchmark for 16–24 GB via Ollama.
ollama run qwen3-coder:30b
The open-weight code reference from late 2024, now surpassed by the Qwen3-Coder / Devstral generation.
ollama run qwen2.5-coder:32b
Coding 14B. HumanEval 89.6, LiveCodeBench 37.1. VRAM sweet spot for self-hosted coding.
ollama run qwen2.5-coder:14b
Dense multimodal 27B released April 22, 2026. 262k ctx (1M YaRN). SWE-bench Verified 77.2%.
ollama run qwen3.6:27b
| Rank | Model | Params | Q4 VRAM | Context | License |
|---|---|---|---|---|---|
| #1 | Laguna XS.2 | 33B | 19 GB | 131 072 | Apache 2.0 |
| #2 | Devstral Small 2 24B | 24B | 14 GB | 256 000 | Apache 2.0 |
| #3 | Laguna XS 2.1 | 33B | 19 GB | 256 000 | OpenMDW 1.1 |
| #4 | Qwen3-Coder 30B-A3B | 30B | 19 GB | 262 144 | Apache 2.0 |
| #5 | Qwen 2.5 Coder 32B | 32B | 19 GB | 131 072 | Apache 2.0 |
| #6 | Qwen 2.5 Coder 14B Instruct | 14B | 9 GB | 131 072 | Apache 2.0 |
| #7 | Qwen 3.6 27B | 27B | 16 GB | 262 144 | Apache 2.0 |
Two options depending on the memory you need and your software: Radeon RX 9070 XT 16 GB · GMKtec EVO-X2 64 GB / 1 TB (Ryzen AI Max+ 395) — compare AI hardware →
You have your model. Now you need to make it work in your editor: the Local Copilot kit connects it to VS Code with Cline (ch. 7), to the terminal with Aider (ch. 8), and to autocomplete as you type (ch. 9)—using the Modelfile tuned for this specific model (ch. 11).
Free memo
Get the memo VRAM → best coding model → Ollama command (one screen, copy and paste). Then switch to the Copilote Local kit for a setup that actually works.
The Local Copilot kit — the Ollama + Cline + Aider configs are ready to paste, with tuned Modelfiles, troubleshooting, and lifetime online access →No spam. Unsubscribe in 1 click. Your data stays with us (never resold).
Your card → the best coding model to run locally, and the exact Ollama command:
| Your VRAM | Typical GPUs / Macs | Recommended coding model | Command Ollama |
|---|---|---|---|
| 8 GB | RTX 4060 / 3060 · M1-M2 16 GB | Qwen 3.5 9B (Q4, 6.6 GB — 256k context) | ollama run qwen3.5:9b |
| 12 GB | RTX 3060 12 GB / 4070 / 5070 | Qwen 3.5 9B (Q8, 11 GB) or Gemma 4 12B (7.6 GB) | ollama run qwen3.5:9b-q8_0 |
| 16 GB | RTX 5070 Ti / 4080 / 5080 · RX 9070 XT · M4 24 GB | Devstral 24B (Q4, 14 GB) — coding-agent specialist | ollama run devstral:24b |
| 24 GB | RTX 3090 / 4090 · RX 7900 XTX · M4 Pro 48 GB | Qwen 3.8 27B (Q4, 18 GB) — the “close to Copilot” option | ollama run qwen3.8:27b |
| 32 GB | RTX 5090 | Qwen 3.6 35B-A3B (Q4, 23 GB) — fast MoE | ollama run qwen3.6:35b |
| 48 GB+ | Mac M4 Max 64 GB · M2 Ultra 128 GB | Qwen3-Coder 30B-A3B (Q8, 32 GB — 256k context) | ollama run qwen3-coder:30b-a3b-q8_0 |
-base : ollama run qwen2.5-coder:7b-base — it’s still the reference for this specific use case. ⚠️ Qwen 3.8: its reasoning is set very high by default and it “overthinks” simple requests — lower it to low (or turn it off) on first launch. ⚠️ License trap: Codestral 22B = Mistral Non-Production License → prohibited for coding at work. Qwen 3.5/3.8, Gemma 4, and Devstral are Apache 2.0. 💡 Running out of memory? Keep ~1.5 GB of VRAM free for context, or drop down one quantization level.🔌 To connect it to VS Code: Cline (multi-file agent), Aider (CLI) or Tabby/Twinny (FIM autocomplete) — they all connect to Ollama locally. The kit Local Copilot — ready-to-paste configs + tested setup — is available: /copilote-local.
We filter models whose tags include “code” and that remain ≤ 100B (beyond that, they’re no longer “local”). The ranking combines code specialization (name contains “coder,” “codestral,” “devstral,” or “laguna”), long context, a permissive license, and a 7–32B sweet-spot size.
Criteria considered:
The scoring is fully transparent: see our methodology for details on VRAM/tokens/sec calculations.
What is the best open-source LLM for coding in 2026?
Our #1 is Laguna XS.2 (33B, Apache 2.0). It combines high HumanEval accuracy, a 131,072-token context, and a permissive license.
How much VRAM do you need for a good coding LLM?
For a 7B Coder model in Q4_K_M, 6–8 GB of VRAM is enough. For a 32B Coder in Q4, plan for 20–24 GB (RTX 4090, 3090, 7900 XTX). Q5 or Q8 climbs quickly — consult the configurator.
Can I use it in commercial production?
Yes, if the license is Apache 2.0 or MIT. Codestral is under Mistral Non-Production License — for personal/research use only. Qwen and DeepSeek are free.
How do you integrate a local LLM into VS Code?
Via the extension Cline, which connects to your local Ollama or LM Studio server. Aider works in the CLI for multi-file refactoring.