Coding

SWE-bench Verified Leaderboard

SWE-bench Verified tests whether a model can resolve real GitHub issues from popular open-source Python projects, scored on the official swebench.com leaderboard as the percentage of human-validated issues actually fixed. It is the headline measure of practical, agentic software-engineering ability — where open-weight models like Qwen3-Coder, GLM-4.6, Kimi K2 and DeepSWE are now competitive with the frontier.

Source: swebench16 open models ranked+146 proprietaryData through Feb 2026

Open models ranked on SWE-bench Verified

# shows rank among open models / rank overall (including proprietary).

#ModelScore
1 / 10MiniMax M2.5 · 228.7B
75.8%
2 / 25GLM 5 · 753.9B
72.8%
3 / 33Kimi K2 Instruct 0905 · 1026.5B
71.2%
4 / 36Kimi K2.5 · 1026.9B
70.8%
5 / 44DeepSeek V3.2 · 685.4B
70.0%
6 / 47Qwen3 Coder 480B A35B Instruct · 480.2B
69.6%
7 / 49GLM 4.6 · 356.8B
68.2%
8 / 60Kimi K2 Instruct · 1026.4B
65.4%
9 / 65GLM 4.5 · 358.3B
64.2%
10 / 67Kimi K2 Thinking · 1026.4B
63.4%
11 / 73MiniMax M2 · 228.7B
61.0%
12 / 75Qwen3 Coder 30B A3B Instruct · 30.5B
60.4%
13 / 78DeepSWE Preview · 32.8B
58.8%
14 / 104Qwen2.5 Coder 32B Instruct · 32.8B
47.0%
15 / 116DeepSeek v3 0324 · 684.5B
42.0%
16 / 142GPT OSS 120B · 116.8B
26.0%

Score vs model size

Which models give the most quality for their size — the ones worth running locally.

100B1Tmodel size (log scale) →75.8%26.0%GLM 5 · 754B · 72.8%Kimi K2 Instruct 0905 · 1T · 71.2%Kimi K2.5 · 1T · 70.8%DeepSeek V3.2 · 685B · 70.0%Qwen3 Coder 480B A35B Instruct · 480B · 69.6%GLM 4.6 · 357B · 68.2%Kimi K2 Instruct · 1T · 65.4%GLM 4.5 · 358B · 64.2%Kimi K2 Thinking · 1T · 63.4%DeepSWE Preview · 33B · 58.8%Qwen2.5 Coder 32B Instruct · 33B · 47.0%DeepSeek v3 0324 · 685B · 42.0%GPT OSS 120B · 117B · 26.0%Qwen3 Coder 30B A3B Instruct · 31B · 60.4%Qwen3 Coder 30B A3B I…MiniMax M2 · 229B · 61.0%MiniMax M2MiniMax M2.5 · 229B · 75.8%MiniMax M2.5
Each dot is a model. Up = higher score, left = smaller (easier to run locally). The dashed line marks the efficiency frontier — the best score you can get at each size or smaller.
  • Qwen3 Coder 30B A3B Instruct, 31B, score 60.4% — on the efficiency frontier (best score at its size or smaller).
  • MiniMax M2, 229B, score 61.0% — on the efficiency frontier (best score at its size or smaller).
  • MiniMax M2.5, 229B, score 75.8% — on the efficiency frontier (best score at its size or smaller).

SWE-bench Verified: frequently asked questions

What is the best open LLM on SWE-bench Verified?
MiniMax M2.5 is the top open model on SWE-bench Verified, scoring 75.8%. Among all models tested — including proprietary ones — it ranks #9. The top model overall is Sonar Foundation Agent + Claude 4.5 Opus at 79.2%.
What's the best SWE-bench Verified model you can run on a 24 GB GPU?
Qwen3 Coder 30B A3B Instruct is the highest-scoring open model that fits in 24 GB at 4-bit quantization (about 17 GB), scoring 60.4% on SWE-bench Verified.
Can open models match proprietary models on SWE-bench Verified?
Not quite on SWE-bench Verified: the strongest proprietary model (Sonar Foundation Agent + Claude 4.5 Opus) scores 79.2%, ahead of the best open model (MiniMax M2.5) at 75.8% — but you can run the open one yourself.

Scores aggregated from swebench. llmrun does not run this benchmark — see the source for methodology, or the about benchmarks for what it measures.