Reasoning

CL-bench Leaderboard

CL-bench, built by Tencent Hunyuan and Fudan University's NLP group, tests whether a model can learn new domain knowledge, rule systems and procedures purely from a long provided context — not from pretraining — and apply them correctly, across 1,899 tasks drawn from 500 scenarios. An LLM judge grades each task against dozens of fine-grained rubrics, and a task counts as solved only when every rubric passes.

Source: epoch7 open models ranked+15 proprietaryData through Mar 2026

Open models ranked on CL-bench

# shows rank among open models / rank overall (including proprietary).

#ModelScore
1 / 10Kimi K2.5 · 1026.9B
19.3%
2 / 11GLM 5 · 753.9B
18.7%
3 / 15Kimi K2 Thinking · 1026.4B
17.6%
4 / 16GLM 4.7 · 358.3B
15.9%
5 / 20DeepSeek V3.2 Exp · 685.4B
13.2%
6 / 21DeepSeek V3.2 · 685.4B
12.4%
7 / 22MiniMax M2.5 · 228.7B
11.4%

Score vs model size

Which models give the most quality for their size — the ones worth running locally.

229B1Tmodel size (log scale) →19.3%11.4%Kimi K2 Thinking · 1T · 17.6%DeepSeek V3.2 Exp · 685B · 13.2%DeepSeek V3.2 · 685B · 12.4%MiniMax M2.5 · 229B · 11.4%MiniMax M2.5GLM 4.7 · 358B · 15.9%GLM 4.7GLM 5 · 754B · 18.7%GLM 5Kimi K2.5 · 1T · 19.3%Kimi K2.5
Each dot is a model. Up = higher score, left = smaller (easier to run locally). The dashed line marks the efficiency frontier — the best score you can get at each size or smaller.
  • MiniMax M2.5, 229B, score 11.4% — on the efficiency frontier (best score at its size or smaller).
  • GLM 4.7, 358B, score 15.9% — on the efficiency frontier (best score at its size or smaller).
  • GLM 5, 754B, score 18.7% — on the efficiency frontier (best score at its size or smaller).
  • Kimi K2.5, 1T, score 19.3% — on the efficiency frontier (best score at its size or smaller).

CL-bench: frequently asked questions

What is the best open LLM on CL-bench?
Kimi K2.5 is the top open model on CL-bench, scoring 19.3%. Among all models tested — including proprietary ones — it ranks #10. The top model overall is GPT 5.4 (Mar 05, 2026, xhigh) (OpenAI) at 27.9%.
Can open models match proprietary models on CL-bench?
Not quite on CL-bench: the strongest proprietary model (GPT 5.4 (Mar 05, 2026, xhigh)) scores 27.9%, ahead of the best open model (Kimi K2.5) at 19.3% — but you can run the open one yourself.

Scores aggregated from epoch. llmrun does not run this benchmark — see the source for methodology, or the about benchmarks for what it measures.