Coding
FrontierSWE Leaderboard
FrontierSWE, from Proximal, tests coding agents on long-horizon software-engineering tasks spanning implementation, performance optimisation and open-ended research, each run for hours through an agent harness. The score is the average normalised task score across repeated runs of every task.
Source: epoch4 open models ranked+15 proprietaryData through Sep 2026
All models ranked on FrontierSWE
Proprietary / closed models are shown dimmed — you can't run them locally, but they show where the open field stands.
| # | Model | Score |
|---|---|---|
| 1 | GPT 6 Astra Max · proprietary | 65.5% |
| 2 | Claude Opus 5.5 Max · proprietary | 62.3% |
| 3 | Claude Sonnet 5.5 Max · proprietary | 61.9% |
| 4 | Claude Fable 5.1 Max · proprietary | 56.3% |
| 5 | Gemini 4 Argon (high) · proprietary | 55.0% |
| 6 | Claude Opus 5 Max · proprietary | 52.0% |
| 7 | Claude Fable 5 Max · proprietary | 47.0% |
| 8 | GPT 5.6 Sol Max · proprietary | 32.2% |
| 9 | GLM 5.3 · 753.3B | 30.2% |
| 10 | Grok 4.7 (xhigh) · proprietary | 29.5% |
| 11 | Kimi K3 · 2779.9B | 25.9% |
| 12 | Grok 4.6 (xhigh) · proprietary | 25.3% |
| 13 | Gemini 3.7 Flash (high) · proprietary | 20.3% |
| 14 | Gemini 3.8 Flash (high) · proprietary | 19.6% |
| 15 | GLM 5.3 Flash · 321.3B | 18.1% |
| 16 | Qwen3.8 Max (Sep 02, xhigh) · proprietary | 17.8% |
| 17 | Qwen3.8 Max (xhigh) · proprietary | 15.8% |
| 18 | Muse Spark 1.2 (xhigh) · proprietary | 12.0% |
| 19 | Inkling · 952.4B | 4.1% |
FrontierSWE: frequently asked questions
- What is the best open LLM on FrontierSWE?
- GLM 5.3 is the top open model on FrontierSWE, scoring 30.2%. Among all models tested — including proprietary ones — it ranks #9. The top model overall is GPT 6 Astra Max (OpenAI) at 65.5%.
- Can open models match proprietary models on FrontierSWE?
- Not quite on FrontierSWE: the strongest proprietary model (GPT 6 Astra Max) scores 65.5%, ahead of the best open model (GLM 5.3) at 30.2% — but you can run the open one yourself.
Scores aggregated from epoch. llmrun does not run this benchmark — see the source for methodology, or the about benchmarks for what it measures.