Coding

FrontierSWE Leaderboard

FrontierSWE, from Proximal, tests coding agents on long-horizon software-engineering tasks spanning implementation, performance optimisation and open-ended research, each run for hours through an agent harness. The score is the average normalised task score across repeated runs of every task.

Source: epoch4 open models ranked+15 proprietaryData through Sep 2026

Open models ranked on FrontierSWE

# shows rank among open models / rank overall (including proprietary).

#ModelScore
1 / 9GLM 5.3 · 753.3B
30.2%
2 / 11Kimi K3 · 2779.9B
25.9%
3 / 15GLM 5.3 Flash · 321.3B
18.1%
4 / 19Inkling · 952.4B
4.1%

FrontierSWE: frequently asked questions

What is the best open LLM on FrontierSWE?
GLM 5.3 is the top open model on FrontierSWE, scoring 30.2%. Among all models tested — including proprietary ones — it ranks #9. The top model overall is GPT 6 Astra Max (OpenAI) at 65.5%.
Can open models match proprietary models on FrontierSWE?
Not quite on FrontierSWE: the strongest proprietary model (GPT 6 Astra Max) scores 65.5%, ahead of the best open model (GLM 5.3) at 30.2% — but you can run the open one yourself.

Scores aggregated from epoch. llmrun does not run this benchmark — see the source for methodology, or the about benchmarks for what it measures.