DeepSWE Leaderboard
DeepSWE, built by Datacurve, evaluates coding agents on 113 original software-engineering tasks across 91 repositories and five languages — written from scratch rather than adapted from real commits, so a model can't have seen the solution during pretraining, and graded Pass@1 with hand-written behavioral verifiers. It's a benchmark name only, unrelated to the open-weight DeepSWE-Preview coding model.
Source: epoch5 open models ranked+63 proprietaryData through Sep 2026
All models ranked on DeepSWE
Proprietary / closed models are shown dimmed — you can't run them locally, but they show where the open field stands.
| # | Model | Score |
|---|---|---|
| 1 | GPT 6 Astra (xhigh) · proprietary | 74.1% |
| 2 | Gemini 3.8 Flash (high) · proprietary | 73.8% |
| 3 | Claude Opus 5 Max · proprietary | 73.7% |
| 4 | GPT 6 Astra (high) · proprietary | 73.2% |
| 5 | GPT 6 Astra Max · proprietary | 73.2% |
| 6 | Claude Opus 5 (xhigh) · proprietary | 73.2% |
| 7 | Claude Opus 5 (high) · proprietary | 72.8% |
| 8 | GPT 6 Astra (medium) · proprietary | 72.8% |
| 9 | GPT 5.6 Sol Max · proprietary | 72.7% |
| 10 | Gemini 3.8 Flash (medium) · proprietary | 71.0% |
| 11 | GPT 5.6 Sol (xhigh) · proprietary | 70.7% |
| 12 | Claude Fable 5 (xhigh) · proprietary | 69.9% |
| 13 | Claude Fable 5 Max · proprietary | 69.7% |
| 14 | GPT 5.6 Terra Max · proprietary | 69.6% |
| 15 | GPT 5.6 Sol (high) · proprietary | 69.4% |
| 16 | GLM 5.3 · 753.3B | 69.0% |
| 17 | Claude Opus 5 (medium) · proprietary | 68.9% |
| 18 | Claude Fable 5 (high) · proprietary | 68.6% |
| 19 | Kimi K3 · 2779.9B | 68.5% |
| 20 | Grok 4.6 (medium) · proprietary | 67.5% |
| 21 | GPT 5.6 Luna Max · proprietary | 67.2% |
| 22 | GPT 5.5 (xhigh) · proprietary | 67.0% |
| 23 | GPT 6 Astra (low) · proprietary | 67.0% |
| 24 | Grok 4.6 (xhigh) · proprietary | 66.7% |
| 25 | Gemini 3.7 Flash (medium) · proprietary | 65.5% |
| 26 | Claude Fable 5 (medium) · proprietary | 65.4% |
| 27 | Gemini 3.7 Flash (high) · proprietary | 65.3% |
| 28 | Grok 4.6 (high) · proprietary | 65.2% |
| 29 | GPT 5.5 (high) · proprietary | 64.4% |
| 30 | GLM 5.3 Flash · 321.3B | 63.4% |
| 31 | GPT 5.6 Sol (medium) · proprietary | 61.1% |
| 32 | GPT 5.6 Terra (xhigh) · proprietary | 60.2% |
| 33 | Claude Fable 5 (low) · proprietary | 59.6% |
| 34 | Claude Opus 4.8 Max · proprietary | 59.0% |
| 35 | Claude Opus 5 (low) · proprietary | 58.1% |
| 36 | Qwen3.8 Max (xhigh) · proprietary | 57.5% |
| 37 | GPT 5.6 Luna (xhigh) · proprietary | 56.9% |
| 38 | Muse Spark 1.2 (xhigh) · proprietary | 54.9% |
| 39 | Claude Opus 4.8 (xhigh) · proprietary | 54.4% |
| 40 | GPT 5.5 (medium) · proprietary | 54.0% |
| 41 | Claude Sonnet 5 Max · proprietary | 53.8% |
| 42 | Gemini 3.7 Flash (low) · proprietary | 53.8% |
| 43 | GPT 5.6 Terra (high) · proprietary | 53.8% |
| 44 | Grok 4.5 (high) · proprietary | 53.8% |
| 45 | Muse Spark 1.1 · proprietary | 53.3% |
| 46 | Claude Opus 4.8 (high) · proprietary | 51.8% |
| 47 | GPT 5.4 (Mar 05, 2026, xhigh) · proprietary | 51.8% |
| 48 | Claude Sonnet 5 (xhigh) · proprietary | 49.7% |
| 49 | Claude Opus 4.8 (medium) · proprietary | 48.7% |
| 50 | Claude Sonnet 5 (high) · proprietary | 48.2% |
| 51 | Gemini 3.6 Flash (high) · proprietary | 46.7% |
| 52 | GPT 5.6 Sol (low) · proprietary | 45.4% |
| 53 | GPT 5.6 Luna (high) · proprietary | 44.3% |
| 54 | GLM 5.2 · 753.3B | 43.8% |
| 55 | Grok 4.6 (low) · proprietary | 41.6% |
| 56 | Claude Opus 4.8 (low) · proprietary | 40.8% |
| 57 | Claude Sonnet 5 (medium) · proprietary | 39.8% |
| 58 | Gemini 3.5 Flash (medium) · proprietary | 37.4% |
| 59 | Gemini 3.5 Flash (high) · proprietary | 36.1% |
| 60 | GPT 5.6 Terra (medium) · proprietary | 35.1% |
| 61 | Kimi K2.7 Code · 1026.9B | 30.5% |
| 62 | Claude Sonnet 5 (low) · proprietary | 30.5% |
| 63 | Claude Sonnet 4.6 (high) · proprietary | 29.9% |
| 64 | GPT 5.5 (low) · proprietary | 27.0% |
| 65 | GPT 5.6 Terra (low) · proprietary | 24.1% |
| 66 | Gemini 3.1 Pro Preview · proprietary | 11.7% |
| 67 | GPT 5.6 Luna (medium) · proprietary | 11.3% |
| 68 | GPT 5.6 Luna (low) · proprietary | 1.6% |
Score vs model size
Which models give the most quality for their size — the ones worth running locally.
- GLM 5.3 Flash, 321B, score 63.4% — on the efficiency frontier (best score at its size or smaller).
- GLM 5.3, 753B, score 69.0% — on the efficiency frontier (best score at its size or smaller).
DeepSWE: frequently asked questions
- What is the best open LLM on DeepSWE?
- GLM 5.3 is the top open model on DeepSWE, scoring 69.0%. Among all models tested — including proprietary ones — it ranks #16. The top model overall is GPT 6 Astra (xhigh) (OpenAI) at 74.1%.
- Can open models match proprietary models on DeepSWE?
- Not quite on DeepSWE: the strongest proprietary model (GPT 6 Astra (xhigh)) scores 74.1%, ahead of the best open model (GLM 5.3) at 69.0% — but you can run the open one yourself.
Scores aggregated from epoch. llmrun does not run this benchmark — see the source for methodology, or the about benchmarks for what it measures.