GSO-Bench Leaderboard
GSO-Bench, built by a team of UC Berkeley-affiliated researchers, gives a coding agent a real codebase and a performance test and asks it to optimize runtime efficiency to match an expert developer's speedup without breaking correctness, across 102 tasks in 10 codebases and five languages. The scored metric, Opt@1, is the share of tasks where a single attempt reaches at least 95% of the human speedup while still passing all tests.
Source: epoch3 open models ranked+35 proprietaryData through Jun 2026
All models ranked on GSO-Bench
Proprietary / closed models are shown dimmed — you can't run them locally, but they show where the open field stands.
| # | Model | Score |
|---|---|---|
| 1 | Claude Opus 4.8 (unspecified) · proprietary | 47.1% |
| 2 | Claude Opus 4.7 · proprietary | 44.1% |
| 3 | Claude Opus 4.7 (high) · proprietary | 44.1% |
| 4 | Claude Opus 4.6 (high) · proprietary | 41.2% |
| 5 | GPT 5.5 (xhigh) · proprietary | 40.2% |
| 6 | Claude Sonnet 5 (unspecified) · proprietary | 37.3% |
| 7 | Claude Opus 4.6 · proprietary | 33.3% |
| 8 | Claude Opus 4.6 (unspecified) · proprietary | 33.3% |
| 9 | GPT 5.4 (Mar 05, 2026, xhigh) · proprietary | 31.4% |
| 10 | GPT 5.2 (Dec 11, 2025, high) · proprietary | 27.5% |
| 11 | Claude Opus 4.5 (Nov 01, 2025) · proprietary | 26.5% |
| 12 | Claude Opus 4.5 (Nov 01, 2025, unspecified) · proprietary | 26.5% |
| 13 | GPT 5.4 (Mar 05, 2026, high) · proprietary | 25.5% |
| 14 | Gemini 3.1 Pro Preview · proprietary | 22.6% |
| 15 | Gemini 3 Pro Preview · proprietary | 18.6% |
| 16 | Claude Sonnet 4.5 (Sep 29, 2025, unspecified) · proprietary | 14.7% |
| 17 | Claude Sonnet 4.5 (Sep 29, 2025) · proprietary | 14.7% |
| 18 | GPT 5.1 (Nov 13, 2025, unspecified) · proprietary | 13.7% |
| 19 | GPT 5.1 (Nov 13, 2025, high) · proprietary | 13.7% |
| 20 | Gemini 3 Flash Preview · proprietary | 9.8% |
| 21 | O3 (Apr 16, 2025, high) · proprietary | 8.8% |
| 22 | Claude Opus 4 (May 14, 2025) · proprietary | 6.9% |
| 23 | Claude Opus 4 (May 14, 2025, unspecified) · proprietary | 6.9% |
| 24 | GPT 5 (Aug 07, 2025, high) · proprietary | 6.9% |
| 25 | Claude Sonnet 4 (May 14, 2025, unspecified) · proprietary | 4.9% |
| 26 | Claude Sonnet 4 (May 14, 2025) · proprietary | 4.9% |
| 27 | Kimi K2 Instruct · 1026.4B | 4.9% |
| 28 | Qwen3 Coder 480B A35B Instruct · 480.2B | 4.9% |
| 29 | Claude 3.5 Sonnet (Oct 22, 2024) · proprietary | 4.6% |
| 30 | Gemini 2.5 Pro · proprietary | 3.9% |
| 31 | Gemini 2.5 Pro Preview (Jun 05) · proprietary | 3.9% |
| 32 | Claude 3.7 Sonnet (Feb 19, 2025, unspecified) · proprietary | 3.8% |
| 33 | Claude 3.7 Sonnet (Feb 19, 2025) · proprietary | 3.8% |
| 34 | O4 Mini (Apr 16, 2025, high) · proprietary | 3.6% |
| 35 | GLM 4.5 Air · 110.5B | 2.9% |
| 36 | O3 Mini (Jan 31, 2025, high) · proprietary | 1.3% |
| 37 | O3 Mini (Jan 31, 2025, low) · proprietary | 1.3% |
| 38 | GPT 4o (Nov 20, 2024) · proprietary | 0.0% |
GSO-Bench: frequently asked questions
- What is the best open LLM on GSO-Bench?
- Kimi K2 Instruct is the top open model on GSO-Bench, scoring 4.9%. Among all models tested — including proprietary ones — it ranks #25. The top model overall is Claude Opus 4.8 (unspecified) (Anthropic) at 47.1%.
- Can open models match proprietary models on GSO-Bench?
- Not quite on GSO-Bench: the strongest proprietary model (Claude Opus 4.8 (unspecified)) scores 47.1%, ahead of the best open model (Kimi K2 Instruct) at 4.9% — but you can run the open one yourself.
Scores aggregated from epoch. llmrun does not run this benchmark — see the source for methodology, or the about benchmarks for what it measures.