Coding

GSO-Bench Leaderboard

GSO-Bench, built by a team of UC Berkeley-affiliated researchers, gives a coding agent a real codebase and a performance test and asks it to optimize runtime efficiency to match an expert developer's speedup without breaking correctness, across 102 tasks in 10 codebases and five languages. The scored metric, Opt@1, is the share of tasks where a single attempt reaches at least 95% of the human speedup while still passing all tests.

Source: epoch3 open models ranked+35 proprietaryData through Jun 2026

All models ranked on GSO-Bench

Proprietary / closed models are shown dimmed — you can't run them locally, but they show where the open field stands.

#ModelScore
1Claude Opus 4.8 (unspecified) · proprietary
47.1%
2Claude Opus 4.7 · proprietary
44.1%
3Claude Opus 4.7 (high) · proprietary
44.1%
4Claude Opus 4.6 (high) · proprietary
41.2%
5GPT 5.5 (xhigh) · proprietary
40.2%
6Claude Sonnet 5 (unspecified) · proprietary
37.3%
7Claude Opus 4.6 · proprietary
33.3%
8Claude Opus 4.6 (unspecified) · proprietary
33.3%
9GPT 5.4 (Mar 05, 2026, xhigh) · proprietary
31.4%
10GPT 5.2 (Dec 11, 2025, high) · proprietary
27.5%
11Claude Opus 4.5 (Nov 01, 2025) · proprietary
26.5%
12Claude Opus 4.5 (Nov 01, 2025, unspecified) · proprietary
26.5%
13GPT 5.4 (Mar 05, 2026, high) · proprietary
25.5%
14Gemini 3.1 Pro Preview · proprietary
22.6%
15Gemini 3 Pro Preview · proprietary
18.6%
16Claude Sonnet 4.5 (Sep 29, 2025, unspecified) · proprietary
14.7%
17Claude Sonnet 4.5 (Sep 29, 2025) · proprietary
14.7%
18GPT 5.1 (Nov 13, 2025, unspecified) · proprietary
13.7%
19GPT 5.1 (Nov 13, 2025, high) · proprietary
13.7%
20Gemini 3 Flash Preview · proprietary
9.8%
21O3 (Apr 16, 2025, high) · proprietary
8.8%
22Claude Opus 4 (May 14, 2025) · proprietary
6.9%
23Claude Opus 4 (May 14, 2025, unspecified) · proprietary
6.9%
24GPT 5 (Aug 07, 2025, high) · proprietary
6.9%
25Claude Sonnet 4 (May 14, 2025, unspecified) · proprietary
4.9%
26Claude Sonnet 4 (May 14, 2025) · proprietary
4.9%
27Kimi K2 Instruct · 1026.4B
4.9%
28Qwen3 Coder 480B A35B Instruct · 480.2B
4.9%
29Claude 3.5 Sonnet (Oct 22, 2024) · proprietary
4.6%
30Gemini 2.5 Pro · proprietary
3.9%
31Gemini 2.5 Pro Preview (Jun 05) · proprietary
3.9%
32Claude 3.7 Sonnet (Feb 19, 2025, unspecified) · proprietary
3.8%
33Claude 3.7 Sonnet (Feb 19, 2025) · proprietary
3.8%
34O4 Mini (Apr 16, 2025, high) · proprietary
3.6%
35GLM 4.5 Air · 110.5B
2.9%
36O3 Mini (Jan 31, 2025, high) · proprietary
1.3%
37O3 Mini (Jan 31, 2025, low) · proprietary
1.3%
38GPT 4o (Nov 20, 2024) · proprietary
0.0%

GSO-Bench: frequently asked questions

What is the best open LLM on GSO-Bench?
Kimi K2 Instruct is the top open model on GSO-Bench, scoring 4.9%. Among all models tested — including proprietary ones — it ranks #25. The top model overall is Claude Opus 4.8 (unspecified) (Anthropic) at 47.1%.
Can open models match proprietary models on GSO-Bench?
Not quite on GSO-Bench: the strongest proprietary model (Claude Opus 4.8 (unspecified)) scores 47.1%, ahead of the best open model (Kimi K2 Instruct) at 4.9% — but you can run the open one yourself.

Scores aggregated from epoch. llmrun does not run this benchmark — see the source for methodology, or the about benchmarks for what it measures.