Coding

DeepSWE Leaderboard

DeepSWE, built by Datacurve, evaluates coding agents on 113 original software-engineering tasks across 91 repositories and five languages — written from scratch rather than adapted from real commits, so a model can't have seen the solution during pretraining, and graded Pass@1 with hand-written behavioral verifiers. It's a benchmark name only, unrelated to the open-weight DeepSWE-Preview coding model.

Source: epoch5 open models ranked+63 proprietaryData through Sep 2026

All models ranked on DeepSWE

Proprietary / closed models are shown dimmed — you can't run them locally, but they show where the open field stands.

#ModelScore
1GPT 6 Astra (xhigh) · proprietary
74.1%
2Gemini 3.8 Flash (high) · proprietary
73.8%
3Claude Opus 5 Max · proprietary
73.7%
4GPT 6 Astra (high) · proprietary
73.2%
5GPT 6 Astra Max · proprietary
73.2%
6Claude Opus 5 (xhigh) · proprietary
73.2%
7Claude Opus 5 (high) · proprietary
72.8%
8GPT 6 Astra (medium) · proprietary
72.8%
9GPT 5.6 Sol Max · proprietary
72.7%
10Gemini 3.8 Flash (medium) · proprietary
71.0%
11GPT 5.6 Sol (xhigh) · proprietary
70.7%
12Claude Fable 5 (xhigh) · proprietary
69.9%
13Claude Fable 5 Max · proprietary
69.7%
14GPT 5.6 Terra Max · proprietary
69.6%
15GPT 5.6 Sol (high) · proprietary
69.4%
16GLM 5.3 · 753.3B
69.0%
17Claude Opus 5 (medium) · proprietary
68.9%
18Claude Fable 5 (high) · proprietary
68.6%
19Kimi K3 · 2779.9B
68.5%
20Grok 4.6 (medium) · proprietary
67.5%
21GPT 5.6 Luna Max · proprietary
67.2%
22GPT 5.5 (xhigh) · proprietary
67.0%
23GPT 6 Astra (low) · proprietary
67.0%
24Grok 4.6 (xhigh) · proprietary
66.7%
25Gemini 3.7 Flash (medium) · proprietary
65.5%
26Claude Fable 5 (medium) · proprietary
65.4%
27Gemini 3.7 Flash (high) · proprietary
65.3%
28Grok 4.6 (high) · proprietary
65.2%
29GPT 5.5 (high) · proprietary
64.4%
30GLM 5.3 Flash · 321.3B
63.4%
31GPT 5.6 Sol (medium) · proprietary
61.1%
32GPT 5.6 Terra (xhigh) · proprietary
60.2%
33Claude Fable 5 (low) · proprietary
59.6%
34Claude Opus 4.8 Max · proprietary
59.0%
35Claude Opus 5 (low) · proprietary
58.1%
36Qwen3.8 Max (xhigh) · proprietary
57.5%
37GPT 5.6 Luna (xhigh) · proprietary
56.9%
38Muse Spark 1.2 (xhigh) · proprietary
54.9%
39Claude Opus 4.8 (xhigh) · proprietary
54.4%
40GPT 5.5 (medium) · proprietary
54.0%
41Claude Sonnet 5 Max · proprietary
53.8%
42Gemini 3.7 Flash (low) · proprietary
53.8%
43GPT 5.6 Terra (high) · proprietary
53.8%
44Grok 4.5 (high) · proprietary
53.8%
45Muse Spark 1.1 · proprietary
53.3%
46Claude Opus 4.8 (high) · proprietary
51.8%
47GPT 5.4 (Mar 05, 2026, xhigh) · proprietary
51.8%
48Claude Sonnet 5 (xhigh) · proprietary
49.7%
49Claude Opus 4.8 (medium) · proprietary
48.7%
50Claude Sonnet 5 (high) · proprietary
48.2%
51Gemini 3.6 Flash (high) · proprietary
46.7%
52GPT 5.6 Sol (low) · proprietary
45.4%
53GPT 5.6 Luna (high) · proprietary
44.3%
54GLM 5.2 · 753.3B
43.8%
55Grok 4.6 (low) · proprietary
41.6%
56Claude Opus 4.8 (low) · proprietary
40.8%
57Claude Sonnet 5 (medium) · proprietary
39.8%
58Gemini 3.5 Flash (medium) · proprietary
37.4%
59Gemini 3.5 Flash (high) · proprietary
36.1%
60GPT 5.6 Terra (medium) · proprietary
35.1%
61Kimi K2.7 Code · 1026.9B
30.5%
62Claude Sonnet 5 (low) · proprietary
30.5%
63Claude Sonnet 4.6 (high) · proprietary
29.9%
64GPT 5.5 (low) · proprietary
27.0%
65GPT 5.6 Terra (low) · proprietary
24.1%
66Gemini 3.1 Pro Preview · proprietary
11.7%
67GPT 5.6 Luna (medium) · proprietary
11.3%
68GPT 5.6 Luna (low) · proprietary
1.6%

Score vs model size

Which models give the most quality for their size — the ones worth running locally.

321B2.8Tmodel size (log scale) →69.0%30.5%Kimi K3 · 2.8T · 68.5%GLM 5.2 · 753B · 43.8%Kimi K2.7 Code · 1T · 30.5%GLM 5.3 Flash · 321B · 63.4%GLM 5.3 FlashGLM 5.3 · 753B · 69.0%GLM 5.3
Each dot is a model. Up = higher score, left = smaller (easier to run locally). The dashed line marks the efficiency frontier — the best score you can get at each size or smaller.
  • GLM 5.3 Flash, 321B, score 63.4% — on the efficiency frontier (best score at its size or smaller).
  • GLM 5.3, 753B, score 69.0% — on the efficiency frontier (best score at its size or smaller).

DeepSWE: frequently asked questions

What is the best open LLM on DeepSWE?
GLM 5.3 is the top open model on DeepSWE, scoring 69.0%. Among all models tested — including proprietary ones — it ranks #16. The top model overall is GPT 6 Astra (xhigh) (OpenAI) at 74.1%.
Can open models match proprietary models on DeepSWE?
Not quite on DeepSWE: the strongest proprietary model (GPT 6 Astra (xhigh)) scores 74.1%, ahead of the best open model (GLM 5.3) at 69.0% — but you can run the open one yourself.

Scores aggregated from epoch. llmrun does not run this benchmark — see the source for methodology, or the about benchmarks for what it measures.