Coding
FrontierSWE Leaderboard
FrontierSWE, from Proximal, tests coding agents on long-horizon software-engineering tasks spanning implementation, performance optimisation and open-ended research, each run for hours through an agent harness. The score is the average normalised task score across repeated runs of every task.
Source: epoch4 open models ranked+15 proprietaryData through Sep 2026
Open models ranked on FrontierSWE
# shows rank among open models / rank overall (including proprietary).
| # | Model | Score |
|---|---|---|
| 1 / 9 | GLM 5.3 · 753.3B | 30.2% |
| 2 / 11 | Kimi K3 · 2779.9B | 25.9% |
| 3 / 15 | GLM 5.3 Flash · 321.3B | 18.1% |
| 4 / 19 | Inkling · 952.4B | 4.1% |
FrontierSWE: frequently asked questions
- What is the best open LLM on FrontierSWE?
- GLM 5.3 is the top open model on FrontierSWE, scoring 30.2%. Among all models tested — including proprietary ones — it ranks #9. The top model overall is GPT 6 Astra Max (OpenAI) at 65.5%.
- Can open models match proprietary models on FrontierSWE?
- Not quite on FrontierSWE: the strongest proprietary model (GPT 6 Astra Max) scores 65.5%, ahead of the best open model (GLM 5.3) at 30.2% — but you can run the open one yourself.
Scores aggregated from epoch. llmrun does not run this benchmark — see the source for methodology, or the about benchmarks for what it measures.