Coding
GSO-Bench Leaderboard
GSO-Bench, built by a team of UC Berkeley-affiliated researchers, gives a coding agent a real codebase and a performance test and asks it to optimize runtime efficiency to match an expert developer's speedup without breaking correctness, across 102 tasks in 10 codebases and five languages. The scored metric, Opt@1, is the share of tasks where a single attempt reaches at least 95% of the human speedup while still passing all tests.
Source: epoch3 open models ranked+35 proprietaryData through Jun 2026
Open models ranked on GSO-Bench
# shows rank among open models / rank overall (including proprietary).
| # | Model | Score |
|---|---|---|
| 1 / 27 | Kimi K2 Instruct · 1026.4B | 4.9% |
| 2 / 28 | Qwen3 Coder 480B A35B Instruct · 480.2B | 4.9% |
| 3 / 35 | GLM 4.5 Air · 110.5B | 2.9% |
GSO-Bench: frequently asked questions
- What is the best open LLM on GSO-Bench?
- Kimi K2 Instruct is the top open model on GSO-Bench, scoring 4.9%. Among all models tested — including proprietary ones — it ranks #25. The top model overall is Claude Opus 4.8 (unspecified) (Anthropic) at 47.1%.
- Can open models match proprietary models on GSO-Bench?
- Not quite on GSO-Bench: the strongest proprietary model (Claude Opus 4.8 (unspecified)) scores 47.1%, ahead of the best open model (Kimi K2 Instruct) at 4.9% — but you can run the open one yourself.
Scores aggregated from epoch. llmrun does not run this benchmark — see the source for methodology, or the about benchmarks for what it measures.