Coding
CursorBench Leaderboard
CursorBench, from Cursor (Anysphere), measures agentic coding on tasks drawn from real Cursor engineering sessions rather than synthetic puzzles. Each model is run at several reasoning-effort levels, and Epoch AI reports the correctness score.
Source: epoch2 open models ranked+57 proprietaryData through Sep 2026
Open models ranked on CursorBench
# shows rank among open models / rank overall (including proprietary).
| # | Model | Score |
|---|---|---|
| 1 / 20 | GLM 5.3 · 753.3B | 42.6% |
| 2 / 33 | GLM 5.3 Flash · 321.3B | 36.8% |
CursorBench: frequently asked questions
- What is the best open LLM on CursorBench?
- GLM 5.3 is the top open model on CursorBench, scoring 42.6%. Among all models tested — including proprietary ones — it ranks #20. The top model overall is Claude Opus 5.5 Max (Anthropic) at 57.8%.
- Can open models match proprietary models on CursorBench?
- Not quite on CursorBench: the strongest proprietary model (Claude Opus 5.5 Max) scores 57.8%, ahead of the best open model (GLM 5.3) at 42.6% — but you can run the open one yourself.
Scores aggregated from epoch. llmrun does not run this benchmark — see the source for methodology, or the about benchmarks for what it measures.