Coding

Terminal-Bench Leaderboard

Terminal-Bench measures whether a model can complete real, end-to-end tasks in a command-line environment — running commands, editing files, and chaining steps — making it an agentic test of practical software skill.

Source: epoch16 open models ranked+41 proprietaryData through Apr 2026

All models ranked on Terminal-Bench

Proprietary / closed models are shown dimmed — you can't run them locally, but they show where the open field stands.

#ModelScore
1Claude Opus 4.7 (unspecified) · proprietary
90.2%
2GPT 5.5 (unspecified) · proprietary
84.7%
3GPT 5.4 (Mar 05, 2026, unspecified) · proprietary
81.8%
4Gemini 3.1 Pro Preview · proprietary
80.2%
5Claude Opus 4.6 (unspecified) · proprietary
79.8%
6GPT 5.3 Codex · proprietary
78.4%
7Claude Opus 4.6 · proprietary
69.9%
8Gemini 3 Pro Preview · proprietary
69.4%
9GPT 5.2 Codex · proprietary
66.5%
10GPT 5.2 (Dec 11, 2025, medium) · proprietary
64.9%
11GPT 5.2 (Dec 11, 2025, unspecified) · proprietary
64.9%
12Gemini 3 Flash Preview · proprietary
64.3%
13Claude Opus 4.5 (Nov 01, 2025, unspecified) · proprietary
63.1%
14Claude Opus 4.5 (Nov 01, 2025) · proprietary
63.1%
15GPT 5.1 Codex Mini · proprietary
61.6%
16GPT 5.1 Codex Max · proprietary
60.4%
17Claude Opus 4.5 (Nov 01, 2025, 128K) · proprietary
59.1%
18GPT 5.1 Codex · proprietary
57.8%
19Grok 4.20 · proprietary
57.3%
20Claude Sonnet 4.6 · proprietary
53.4%
21Claude Sonnet 4.6 (unspecified) · proprietary
53.4%
22GLM 5 · 753.9B
52.4%
23GPT 5 (Aug 07, 2025, medium) · proprietary
49.6%
24GPT 5 (Aug 07, 2025, unspecified) · proprietary
49.6%
25GPT 5.1 (Nov 13, 2025, medium) · proprietary
47.6%
26GPT 5.1 (Nov 13, 2025, unspecified) · proprietary
47.6%
27Claude Sonnet 4.5 (Sep 29, 2025, unspecified) · proprietary
46.5%
28MiniMax M2.7 · 228.7B
45.1%
29GPT 5 Codex · proprietary
44.3%
30Kimi K2.5 · 1058.6B
43.2%
31Claude Sonnet 4.5 (Sep 29, 2025) · proprietary
42.8%
32MiniMax M2.5 · 228.7B
42.7%
33DeepSeek V3.2 · 685.4B
39.6%
34Claude Opus 4.1 (Aug 05, 2025, unspecified) · proprietary
38.0%
35Claude Opus 4.1 (Aug 05, 2025) · proprietary
38.0%
36MiniMax M2.1 · 228.7B
36.6%
37Kimi K2 Thinking · 1058.1B
35.7%
38Claude Haiku 4.5 (Oct 01, 2025, unspecified) · proprietary
35.5%
39GPT 5 Mini (Aug 07, 2025, unspecified) · proprietary
34.8%
40GLM 4.7 · 358.3B
33.4%
41Gemini 2.5 Pro · proprietary
32.6%
42GPT 5 Mini (Aug 07, 2025, medium) · proprietary
31.9%
43MiniMax M2 · 228.7B
30.0%
44Claude Haiku 4.5 (Oct 01, 2025) · proprietary
29.8%
45Kimi K2 Instruct · 1026.5B
27.8%
46Grok 4 (Jul 09) · proprietary
27.2%
47Qwen3 Coder 480B A35B Instruct · 480.2B
27.2%
48Grok Code Fast 1 · proprietary
25.8%
49Qwen3.6 35B A3B · 36.0B
24.6%
50GLM 4.6 · 356.8B
24.5%
51GPT 5 Nano (Aug 07, 2025, unspecified) · proprietary
21.8%
52GPT OSS 120B · 120.4B
18.7%
53Gemini 2.5 Flash · proprietary
17.1%
54Gemini 2.5 Flash Preview (Sep 2025) · proprietary
17.1%
55GPT 5 Nano (Aug 07, 2025, medium) · proprietary
11.5%
56Qwen3.5 9B · 9.7B
9.2%
57GPT OSS 20B · 21.5B
3.4%

Score vs model size

Which models give the most quality for their size — the ones worth running locally.

10B100B1Tmodel size (log scale) →52.4%3.4%Kimi K2.5 · 1.1T · 43.2%MiniMax M2.5 · 229B · 42.7%DeepSeek V3.2 · 685B · 39.6%MiniMax M2.1 · 229B · 36.6%Kimi K2 Thinking · 1.1T · 35.7%GLM 4.7 · 358B · 33.4%MiniMax M2 · 229B · 30.0%Kimi K2 Instruct · 1T · 27.8%Qwen3 Coder 480B A35B Instruct · 480B · 27.2%GLM 4.6 · 357B · 24.5%GPT OSS 120B · 120B · 18.7%GPT OSS 20B · 22B · 3.4%Qwen3.5 9B · 10B · 9.2%Qwen3.5 9BQwen3.6 35B A3B · 36B · 24.6%Qwen3.6 35B A3BMiniMax M2.7 · 229B · 45.1%MiniMax M2.7GLM 5 · 754B · 52.4%GLM 5
Each dot is a model. Up = higher score, left = smaller (easier to run locally). The dashed line marks the efficiency frontier — the best score you can get at each size or smaller.
  • Qwen3.5 9B, 10B, score 9.2% — on the efficiency frontier (best score at its size or smaller).
  • Qwen3.6 35B A3B, 36B, score 24.6% — on the efficiency frontier (best score at its size or smaller).
  • MiniMax M2.7, 229B, score 45.1% — on the efficiency frontier (best score at its size or smaller).
  • GLM 5, 754B, score 52.4% — on the efficiency frontier (best score at its size or smaller).

Terminal-Bench: frequently asked questions

What is the best open LLM on Terminal-Bench?
GLM 5 is the top open model on Terminal-Bench, scoring 52.4%. Among all models tested — including proprietary ones — it ranks #22. The top model overall is Claude Opus 4.7 (unspecified) (Anthropic) at 90.2%.
What's the best Terminal-Bench model you can run on a 24 GB GPU?
Qwen3.6 35B A3B is the highest-scoring open model that fits in 24 GB at 4-bit quantization (about 20 GB), scoring 24.6% on Terminal-Bench.
What's the best Terminal-Bench model you can run on a 12 GB GPU?
Qwen3.5 9B is the highest-scoring open model that fits in 12 GB at 4-bit quantization (about 5 GB), scoring 9.2% on Terminal-Bench.
Can open models match proprietary models on Terminal-Bench?
Not quite on Terminal-Bench: the strongest proprietary model (Claude Opus 4.7 (unspecified)) scores 90.2%, ahead of the best open model (GLM 5) at 52.4% — but you can run the open one yourself.

Scores aggregated from epoch. llmrun does not run this benchmark — see the source for methodology, or the about benchmarks for what it measures.