SWE-bench: the 2026 ranking of open-source LLMs for coding
Evaluating an LLM on SWE-bench locally measures an open-weights model's real ability to solve authentic GitHub bugs, far beyond HumanEval scores. SWE-bench stands out because it requires the model to navigate a complete repository, understand an issue ticket, modify multiple files, and pass the test suite. This article explains the benchmark's mechanics, the SWE-bench Verified variant, hardware requirements, candidate models in the catalog, the local execution procedure, and answers frequently asked questions from self-hosting practitioners.
Understanding the mechanics of SWE-bench
SWE-bench is a benchmark published in 2023 by Princeton University that collects 2,294 real GitHub issues from 12 popular Python projects (Django, scikit-learn, sympy, matplotlib, etc.). Each task gives the model a repository at a specific commit and the issue description. The LLM must produce a patch ("diff) which, once applied, makes the previously failing tests pass without breaking existing tests. For details on dataset construction, see the original SWE-bench paper and the official GitHub repository.
Unlike HumanEval, which evaluates isolated short functions, SWE-bench measures:
- Contextual understanding : read a repository with tens of thousands of lines
- Bug localization : identify the right files to modify without explicit guidance
- Multi-file editing : produce a coherent patch that breaks nothing
- Long-chain reasoning : alternate exploration, hypothesis, and verification
A raw score of 20% on SWE-bench often corresponds to a score above 80% on HumanEval, which explains why many models post excellent HumanEval results but collapse on SWE-bench.
SWE-bench Verified: the filtered version
SWE-bench Verified is a subset of 500 instances manually annotated by OpenAI in collaboration with the Princeton authors. The goal: remove tasks whose problem statement is ambiguous, whose hidden tests are too specific, or whose reference patch depends on context that was not provided. Full details on the OpenAI blog dedicated to Verified.
For a local test, SWE-bench Verified is the right choice :
- 500 instances instead of 2,294: execution feasible in a few days on a workstation
- More reliable evaluation of the model's actual reasoning
- Direct comparison with scores published by vendors
The most widely used reference agent is SWE-agent, a harness that exposes an interactive terminal to the LLM, allowing it to edit files, run commands, and execute tests. A popular alternative is OpenHands, which natively supports OpenAI-compatible backends (vLLM, llama.cpp server, SGLang).
Catalog models suited to the benchmark
The agentic coding LLM must combine a large context window (the harness injects stack traces, source code, and action history) with strong reasoning performance. Here are the catalog candidates sorted by hardware profile.
Very high-VRAM workstation stations (≥ 400 GB)
- DeepSeek V3.2 (685B, MIT)—Q4 VRAM ~410 GB, 128k context. The current open-weight reference on SWE-bench Verified according to independent evaluations.
- DeepSeek R1 671B (671B, MIT) — Q4 VRAM ~400 GB, 128k ctx. Long-chain reasoning model, particularly well suited to bug localization.
- Mistral Large 3 675B (675B, Apache 2.0) — ~405 GB Q4 VRAM, 256k context. Permissive license, interesting for commercial use.
- Kimi K2.6 (1000B, Modified MIT) — Q4 VRAM ~600 GB, 256k context. Optimized for agentic workflows according to Moonshot.
Multi-GPU cluster (140–250 GB)
- Qwen 3 235B-A22B (235B, Apache 2.0) — Q4 VRAM ~142 GB, 131k context. 22B active MoE, good cost-to-quality ratio.
- Llama 4 Maverick 400B (400B, Llama 4 Community)—Q4 VRAM ~240 GB, ctx 1M. Exceptional context for large repositories.
- GLM-5.1 (744B, MIT) — Q4 VRAM ~445 GB, 200k ctx.
Single system (≤ 80 GB)
- Qwen3-Coder-Next 80B-A3B (80B, Apache 2.0) — ~48 GB VRAM in Q4, 262k context. Specialized for code, runnable on an A100 80GB or two 3090s.
- gpt-oss 120B (117B, Apache 2.0)—Q4 VRAM ~70 GB, ctx 128k. An OpenAI model released with open weights.
- Mistral Small 4 (119B, Apache 2.0)—Q4 VRAM ~72 GB, ctx 256k.
For a detailed, real-world code comparison, see Qwen3-Coder vs DeepSeek V3.2 and the page best LLM for coding.
Tokens/sec and impact on run duration
A complete SWE-bench Verified run consumes between 500 million and 2 billion tokens, depending on the harness and the number of allowed attempts. Tokens/sec throughput therefore directly determines the benchmark duration.
Indicative estimates (to be confirmed on your setup) :
- DeepSeek V3.2 on 8×H100 80GB in Q4: ~25 tokens/sec during generation, Verified run estimated at between 5 and 8 days
- Qwen 3 235B-A22B on 4×H100: ~40 tokens/sec in Q4 (22B active MoE helps), estimated run time of 3 to 5 days
- Qwen3-Coder-Next 80B-A3B on 2×A100 80GB: ~60 tokens/sec in Q4, estimated run time of 2 to 3 days
- gpt-oss 120B on 1×H200 141GB in Q4: ~35 tokens/sec, estimated run time 3 to 4 days
Effective context matters as much as raw speed: SWE-agent regularly injects 30k to 60k tokens into the prompt to reconstruct the repository state. A model limited to 32k context is poorly suited; aim for 128k minimum. Also see the VRAM guide by quantization to adjust your memory budget.
Local execution procedure
Here is a simplified procedure for a strict self-hosted setup.
-
Prepare the Docker environment : SWE-bench Verified requires Docker images for each project (Django, sympy, etc.) to isolate test execution. Allow 80 GB of disk space for the official images published by the maintainers.
-
Start an OpenAI-compatible inference server : with vLLM or llama.cpp in server mode, expose your model on
localhost:8000/v1. For Qwen3-Coder-Next in Q4 on 2×A100, a typical vLLM command configures--tensor-parallel-size 2 --max-model-len 131072 --quantization awq. -
Clone SWE-agent and configure the harness to point to the local endpoint. The configuration file accepts
api_base: http://localhost:8000/v1etmodel_name: <votre-modèle>. -
Start a calibration run across 10–20 instances to verify that the action formatting is respected. Many open-weight models fail here because they invent XML tags that the harness does not recognize.
-
Complete run : 500 instances, several days, systematically log the generated patches for postmortem analysis.
-
Optional submission au official leaderboard for public comparison.
Practical tip: limit the number of actions per instance (50–75) to prevent a model from entering an infinite loop that consumes your token budget.
Interpreting the results
A raw score on SWE-bench Verified should be interpreted relatively, not absolutely.
- < 10 % : the model does not understand the agent's action format, or its context is too short
- 10-25 % : adequate level for a general-purpose open-weights model
- 25-50 % : expected level of a specialized reasoning model (DeepSeek R1 671B, Kimi K2.6)
- > 50 % : frontier-level, reached by the best proprietary models
Beyond the score, look at the distribution of failures : tests that time out, patches that do not apply, files modified outside the scope. This analysis often reveals that the model is limited not by its reasoning but by the harness (window size, action parsing). To learn more about selecting an agentic model, see the LLM guide for agents.
FAQ
Q: SWE-bench or SWE-bench Verified for a first test?
Verified, without hesitation. All 500 instances are manually annotated, ambiguities are removed, and runtime remains reasonable on a single workstation. The full SWE-bench (2,294 instances) contains poorly specified tasks that unfairly penalize models. Verified also offers a direct comparison with scores published by proprietary model vendors.
Q: Which quantization should you choose to preserve the code score?
Q4_K_M (GGUF) or 4-bit AWQ typically reduce the SWE-bench score by 1 to 3 points compared with FP16, which remains acceptable. Avoid Q3 and Q2 on reasoning models, where the loss becomes significant. If your VRAM allows it, Q5_K_M or Q8 offer a better trade-off. See the GGUF vs. AWQ quantization comparison.
Q: Do I need a specialized “coder” model or a general-purpose one?
Both profiles work. Code-specialized models such as Qwen3-Coder-Next 80B-A3B excellent at generating patches, but general-purpose reasoning models such as DeepSeek R1 671B are often better at bug localization and multi-step planning, which dominate the cost of a SWE-bench task.
Q: Can SWE-bench Verified run on a single RTX 4090?
Difficult. 24 GB of VRAM limits you to models ≤ 30B in Q4, and these models generally plateau below 10% on Verified due to insufficient reasoning. You can use it to validate your pipeline with a lightweight model such as Qwen3-Coder-Next with partial CPU offload, but aim for 2×3090 or 1×A100 80GB for a usable result.
Q: How many tokens does a complete run consume?
Estimated at between 500 million and 2 billion tokens depending on the harness (SWE-agent, OpenHands, custom) and the maximum number of actions allowed per instance. With a budget of 50 actions per instance and an average context of 30k tokens, expect about 1 billion tokens for 500 instances. The more the model “thinks” in a chain (R1-style), the higher the token bill climbs.
Q: Are local results comparable to leaderboard scores?
Yes, provided you use the official harness, the same dataset version, and follow the submission protocol (one patch per instance, no oracle retries). Any modification to the harness or system prompt must be documented. The SWE-bench leaderboard accepts self-hosted submissions and verifies reproducibility.
Conclusion
Running an SWE-bench LLM locally remains the most revealing test for distinguishing an open-weights coding model from merely a good HumanEval performer. With SWE-bench Verified, a 2×A100 or 4×H100 workstation and a suitable model (Qwen3-Coder-Next, DeepSeek V3.2 or gpt-oss 120B), a realistic run takes several days. To calibrate your setup before launching the benchmark, use the quelllm.fr configurator or browse the full catalog.