Coding

CursorBench Leaderboard

CursorBench, from Cursor (Anysphere), measures agentic coding on tasks drawn from real Cursor engineering sessions rather than synthetic puzzles. Each model is run at several reasoning-effort levels, and Epoch AI reports the correctness score.

Source: epoch2 open models ranked+57 proprietaryData through Sep 2026

Open models ranked on CursorBench

# shows rank among open models / rank overall (including proprietary).

#ModelScore
1 / 20GLM 5.3 · 753.3B
42.6%
2 / 33GLM 5.3 Flash · 321.3B
36.8%

CursorBench: frequently asked questions

What is the best open LLM on CursorBench?
GLM 5.3 is the top open model on CursorBench, scoring 42.6%. Among all models tested — including proprietary ones — it ranks #20. The top model overall is Claude Opus 5.5 Max (Anthropic) at 57.8%.
Can open models match proprietary models on CursorBench?
Not quite on CursorBench: the strongest proprietary model (Claude Opus 5.5 Max) scores 57.8%, ahead of the best open model (GLM 5.3) at 42.6% — but you can run the open one yourself.

Scores aggregated from epoch. llmrun does not run this benchmark — see the source for methodology, or the about benchmarks for what it measures.