Coding

FrontierSWE Leaderboard

FrontierSWE, from Proximal, tests coding agents on long-horizon software-engineering tasks spanning implementation, performance optimisation and open-ended research, each run for hours through an agent harness. The score is the average normalised task score across repeated runs of every task.

Source: epoch4 open models ranked+15 proprietaryData through Sep 2026

All models ranked on FrontierSWE

Proprietary / closed models are shown dimmed — you can't run them locally, but they show where the open field stands.

#ModelScore
1GPT 6 Astra Max · proprietary
65.5%
2Claude Opus 5.5 Max · proprietary
62.3%
3Claude Sonnet 5.5 Max · proprietary
61.9%
4Claude Fable 5.1 Max · proprietary
56.3%
5Gemini 4 Argon (high) · proprietary
55.0%
6Claude Opus 5 Max · proprietary
52.0%
7Claude Fable 5 Max · proprietary
47.0%
8GPT 5.6 Sol Max · proprietary
32.2%
9GLM 5.3 · 753.3B
30.2%
10Grok 4.7 (xhigh) · proprietary
29.5%
11Kimi K3 · 2779.9B
25.9%
12Grok 4.6 (xhigh) · proprietary
25.3%
13Gemini 3.7 Flash (high) · proprietary
20.3%
14Gemini 3.8 Flash (high) · proprietary
19.6%
15GLM 5.3 Flash · 321.3B
18.1%
16Qwen3.8 Max (Sep 02, xhigh) · proprietary
17.8%
17Qwen3.8 Max (xhigh) · proprietary
15.8%
18Muse Spark 1.2 (xhigh) · proprietary
12.0%
19Inkling · 952.4B
4.1%

FrontierSWE: frequently asked questions

What is the best open LLM on FrontierSWE?
GLM 5.3 is the top open model on FrontierSWE, scoring 30.2%. Among all models tested — including proprietary ones — it ranks #9. The top model overall is GPT 6 Astra Max (OpenAI) at 65.5%.
Can open models match proprietary models on FrontierSWE?
Not quite on FrontierSWE: the strongest proprietary model (GPT 6 Astra Max) scores 65.5%, ahead of the best open model (GLM 5.3) at 30.2% — but you can run the open one yourself.

Scores aggregated from epoch. llmrun does not run this benchmark — see the source for methodology, or the about benchmarks for what it measures.