CL-bench Leaderboard
CL-bench, built by Tencent Hunyuan and Fudan University's NLP group, tests whether a model can learn new domain knowledge, rule systems and procedures purely from a long provided context — not from pretraining — and apply them correctly, across 1,899 tasks drawn from 500 scenarios. An LLM judge grades each task against dozens of fine-grained rubrics, and a task counts as solved only when every rubric passes.
Source: epoch7 open models ranked+15 proprietaryData through Mar 2026
All models ranked on CL-bench
Proprietary / closed models are shown dimmed — you can't run them locally, but they show where the open field stands.
| # | Model | Score |
|---|---|---|
| 1 | GPT 5.4 (Mar 05, 2026, xhigh) · proprietary | 27.9% |
| 2 | GPT 5.1 (Nov 13, 2025, high) · proprietary | 23.7% |
| 3 | Grok 4.20 · proprietary | 22.2% |
| 4 | Claude Opus 4.5 (Nov 01, 2025, unspecified) · proprietary | 21.1% |
| 5 | GPT 5.1 (Nov 13, 2025, unspecified) · proprietary | 21.1% |
| 6 | Gemini 3.1 Pro Preview · proprietary | 20.8% |
| 7 | Claude Opus 4.6 (unspecified) · proprietary | 20.7% |
| 8 | Qwen3.6 Plus · proprietary | 20.3% |
| 9 | Qwen3.5 Plus · proprietary | 19.8% |
| 10 | Kimi K2.5 · 1026.9B | 19.3% |
| 11 | GLM 5 · 753.9B | 18.7% |
| 12 | GPT 5.2 (Dec 11, 2025, unspecified) · proprietary | 18.2% |
| 13 | GPT 5.2 (Dec 11, 2025, high) · proprietary | 18.1% |
| 14 | O3 (Apr 16, 2025, high) · proprietary | 17.8% |
| 15 | Kimi K2 Thinking · 1026.4B | 17.6% |
| 16 | GLM 4.7 · 358.3B | 15.9% |
| 17 | Gemini 3 Pro Preview · proprietary | 15.8% |
| 18 | Mimo v2 Pro · proprietary | 15.7% |
| 19 | Qwen3 Max (Sep 23, 2025) · proprietary | 14.5% |
| 20 | DeepSeek V3.2 Exp · 685.4B | 13.2% |
| 21 | DeepSeek V3.2 · 685.4B | 12.4% |
| 22 | MiniMax M2.5 · 228.7B | 11.4% |
Score vs model size
Which models give the most quality for their size — the ones worth running locally.
- MiniMax M2.5, 229B, score 11.4% — on the efficiency frontier (best score at its size or smaller).
- GLM 4.7, 358B, score 15.9% — on the efficiency frontier (best score at its size or smaller).
- GLM 5, 754B, score 18.7% — on the efficiency frontier (best score at its size or smaller).
- Kimi K2.5, 1T, score 19.3% — on the efficiency frontier (best score at its size or smaller).
CL-bench: frequently asked questions
- What is the best open LLM on CL-bench?
- Kimi K2.5 is the top open model on CL-bench, scoring 19.3%. Among all models tested — including proprietary ones — it ranks #10. The top model overall is GPT 5.4 (Mar 05, 2026, xhigh) (OpenAI) at 27.9%.
- Can open models match proprietary models on CL-bench?
- Not quite on CL-bench: the strongest proprietary model (GPT 5.4 (Mar 05, 2026, xhigh)) scores 27.9%, ahead of the best open model (Kimi K2.5) at 19.3% — but you can run the open one yourself.
Scores aggregated from epoch. llmrun does not run this benchmark — see the source for methodology, or the about benchmarks for what it measures.