Reasoning
CL-bench Life Leaderboard
CL-bench Life, a follow-up from the CL-bench authors, tests the same context-learning skill on messier real-life material — fragmented notes, casual conversations, and behavioral logs — across 405 human-curated context-task pairs. It uses the same rubric-based LLM-judge scoring as CL-bench, requiring every rubric on a task to pass for it to count as solved.
Source: epoch6 open models ranked+11 proprietaryData through Apr 2026
Open models ranked on CL-bench Life
# shows rank among open models / rank overall (including proprietary).
| # | Model | Score |
|---|---|---|
| 1 / 9 | DeepSeek V4 Pro · 1598.8B | 13.5% |
| 2 / 10 | Kimi K2.5 · 1026.9B | 13.2% |
| 3 / 13 | GLM 4.7 · 358.3B | 10.9% |
| 4 / 14 | DeepSeek V3.2 Exp · 685.4B | 9.5% |
| 5 / 15 | DeepSeek V3.2 · 685.4B | 7.4% |
| 6 / 17 | MiniMax M2.5 · 228.7B | 6.3% |
Score vs model size
Which models give the most quality for their size — the ones worth running locally.
- MiniMax M2.5, 229B, score 6.3% — on the efficiency frontier (best score at its size or smaller).
- GLM 4.7, 358B, score 10.9% — on the efficiency frontier (best score at its size or smaller).
- Kimi K2.5, 1T, score 13.2% — on the efficiency frontier (best score at its size or smaller).
- DeepSeek V4 Pro, 1.6T, score 13.5% — on the efficiency frontier (best score at its size or smaller).
CL-bench Life: frequently asked questions
- What is the best open LLM on CL-bench Life?
- DeepSeek V4 Pro is the top open model on CL-bench Life, scoring 13.5%. Among all models tested — including proprietary ones — it ranks #9. The top model overall is GPT 5.5 (high) (OpenAI) at 22.2%.
- Can open models match proprietary models on CL-bench Life?
- Not quite on CL-bench Life: the strongest proprietary model (GPT 5.5 (high)) scores 22.2%, ahead of the best open model (DeepSeek V4 Pro) at 13.5% — but you can run the open one yourself.
Scores aggregated from epoch. llmrun does not run this benchmark — see the source for methodology, or the about benchmarks for what it measures.