Reasoning

CL-bench Leaderboard

CL-bench, built by Tencent Hunyuan and Fudan University's NLP group, tests whether a model can learn new domain knowledge, rule systems and procedures purely from a long provided context — not from pretraining — and apply them correctly, across 1,899 tasks drawn from 500 scenarios. An LLM judge grades each task against dozens of fine-grained rubrics, and a task counts as solved only when every rubric passes.

Source: epoch7 open models ranked+15 proprietaryData through Mar 2026

All models ranked on CL-bench

Proprietary / closed models are shown dimmed — you can't run them locally, but they show where the open field stands.

#ModelScore
1GPT 5.4 (Mar 05, 2026, xhigh) · proprietary
27.9%
2GPT 5.1 (Nov 13, 2025, high) · proprietary
23.7%
3Grok 4.20 · proprietary
22.2%
4Claude Opus 4.5 (Nov 01, 2025, unspecified) · proprietary
21.1%
5GPT 5.1 (Nov 13, 2025, unspecified) · proprietary
21.1%
6Gemini 3.1 Pro Preview · proprietary
20.8%
7Claude Opus 4.6 (unspecified) · proprietary
20.7%
8Qwen3.6 Plus · proprietary
20.3%
9Qwen3.5 Plus · proprietary
19.8%
10Kimi K2.5 · 1026.9B
19.3%
11GLM 5 · 753.9B
18.7%
12GPT 5.2 (Dec 11, 2025, unspecified) · proprietary
18.2%
13GPT 5.2 (Dec 11, 2025, high) · proprietary
18.1%
14O3 (Apr 16, 2025, high) · proprietary
17.8%
15Kimi K2 Thinking · 1026.4B
17.6%
16GLM 4.7 · 358.3B
15.9%
17Gemini 3 Pro Preview · proprietary
15.8%
18Mimo v2 Pro · proprietary
15.7%
19Qwen3 Max (Sep 23, 2025) · proprietary
14.5%
20DeepSeek V3.2 Exp · 685.4B
13.2%
21DeepSeek V3.2 · 685.4B
12.4%
22MiniMax M2.5 · 228.7B
11.4%

Score vs model size

Which models give the most quality for their size — the ones worth running locally.

229B1Tmodel size (log scale) →19.3%11.4%Kimi K2 Thinking · 1T · 17.6%DeepSeek V3.2 Exp · 685B · 13.2%DeepSeek V3.2 · 685B · 12.4%MiniMax M2.5 · 229B · 11.4%MiniMax M2.5GLM 4.7 · 358B · 15.9%GLM 4.7GLM 5 · 754B · 18.7%GLM 5Kimi K2.5 · 1T · 19.3%Kimi K2.5
Each dot is a model. Up = higher score, left = smaller (easier to run locally). The dashed line marks the efficiency frontier — the best score you can get at each size or smaller.
  • MiniMax M2.5, 229B, score 11.4% — on the efficiency frontier (best score at its size or smaller).
  • GLM 4.7, 358B, score 15.9% — on the efficiency frontier (best score at its size or smaller).
  • GLM 5, 754B, score 18.7% — on the efficiency frontier (best score at its size or smaller).
  • Kimi K2.5, 1T, score 19.3% — on the efficiency frontier (best score at its size or smaller).

CL-bench: frequently asked questions

What is the best open LLM on CL-bench?
Kimi K2.5 is the top open model on CL-bench, scoring 19.3%. Among all models tested — including proprietary ones — it ranks #10. The top model overall is GPT 5.4 (Mar 05, 2026, xhigh) (OpenAI) at 27.9%.
Can open models match proprietary models on CL-bench?
Not quite on CL-bench: the strongest proprietary model (GPT 5.4 (Mar 05, 2026, xhigh)) scores 27.9%, ahead of the best open model (Kimi K2.5) at 19.3% — but you can run the open one yourself.

Scores aggregated from epoch. llmrun does not run this benchmark — see the source for methodology, or the about benchmarks for what it measures.