CL-bench Leaderboard
CL-bench, built by Tencent Hunyuan and Fudan University's NLP group, tests whether a model can learn new domain knowledge, rule systems and procedures purely from a long provided context — not from pretraining — and apply them correctly, across 1,899 tasks drawn from 500 scenarios. An LLM judge grades each task against dozens of fine-grained rubrics, and a task counts as solved only when every rubric passes.
Source: epoch7 open models ranked+15 proprietaryData through Mar 2026
Open models ranked on CL-bench
# shows rank among open models / rank overall (including proprietary).
| # | Model | Score |
|---|---|---|
| 1 / 10 | Kimi K2.5 · 1026.9B | 19.3% |
| 2 / 11 | GLM 5 · 753.9B | 18.7% |
| 3 / 15 | Kimi K2 Thinking · 1026.4B | 17.6% |
| 4 / 16 | GLM 4.7 · 358.3B | 15.9% |
| 5 / 20 | DeepSeek V3.2 Exp · 685.4B | 13.2% |
| 6 / 21 | DeepSeek V3.2 · 685.4B | 12.4% |
| 7 / 22 | MiniMax M2.5 · 228.7B | 11.4% |
Score vs model size
Which models give the most quality for their size — the ones worth running locally.
- MiniMax M2.5, 229B, score 11.4% — on the efficiency frontier (best score at its size or smaller).
- GLM 4.7, 358B, score 15.9% — on the efficiency frontier (best score at its size or smaller).
- GLM 5, 754B, score 18.7% — on the efficiency frontier (best score at its size or smaller).
- Kimi K2.5, 1T, score 19.3% — on the efficiency frontier (best score at its size or smaller).
CL-bench: frequently asked questions
- What is the best open LLM on CL-bench?
- Kimi K2.5 is the top open model on CL-bench, scoring 19.3%. Among all models tested — including proprietary ones — it ranks #10. The top model overall is GPT 5.4 (Mar 05, 2026, xhigh) (OpenAI) at 27.9%.
- Can open models match proprietary models on CL-bench?
- Not quite on CL-bench: the strongest proprietary model (GPT 5.4 (Mar 05, 2026, xhigh)) scores 27.9%, ahead of the best open model (Kimi K2.5) at 19.3% — but you can run the open one yourself.
Scores aggregated from epoch. llmrun does not run this benchmark — see the source for methodology, or the about benchmarks for what it measures.