Reasoning

CL-bench Life Leaderboard

CL-bench Life, a follow-up from the CL-bench authors, tests the same context-learning skill on messier real-life material — fragmented notes, casual conversations, and behavioral logs — across 405 human-curated context-task pairs. It uses the same rubric-based LLM-judge scoring as CL-bench, requiring every rubric on a task to pass for it to count as solved.

Source: epoch6 open models ranked+11 proprietaryData through Apr 2026

Open models ranked on CL-bench Life

# shows rank among open models / rank overall (including proprietary).

#ModelScore
1 / 9DeepSeek V4 Pro · 1598.8B
13.5%
2 / 10Kimi K2.5 · 1026.9B
13.2%
3 / 13GLM 4.7 · 358.3B
10.9%
4 / 14DeepSeek V3.2 Exp · 685.4B
9.5%
5 / 15DeepSeek V3.2 · 685.4B
7.4%
6 / 17MiniMax M2.5 · 228.7B
6.3%

Score vs model size

Which models give the most quality for their size — the ones worth running locally.

229B1.6Tmodel size (log scale) →13.5%6.3%DeepSeek V3.2 Exp · 685B · 9.5%DeepSeek V3.2 · 685B · 7.4%MiniMax M2.5 · 229B · 6.3%MiniMax M2.5GLM 4.7 · 358B · 10.9%GLM 4.7Kimi K2.5 · 1T · 13.2%Kimi K2.5DeepSeek V4 Pro · 1.6T · 13.5%DeepSeek V4 Pro
Each dot is a model. Up = higher score, left = smaller (easier to run locally). The dashed line marks the efficiency frontier — the best score you can get at each size or smaller.
  • MiniMax M2.5, 229B, score 6.3% — on the efficiency frontier (best score at its size or smaller).
  • GLM 4.7, 358B, score 10.9% — on the efficiency frontier (best score at its size or smaller).
  • Kimi K2.5, 1T, score 13.2% — on the efficiency frontier (best score at its size or smaller).
  • DeepSeek V4 Pro, 1.6T, score 13.5% — on the efficiency frontier (best score at its size or smaller).

CL-bench Life: frequently asked questions

What is the best open LLM on CL-bench Life?
DeepSeek V4 Pro is the top open model on CL-bench Life, scoring 13.5%. Among all models tested — including proprietary ones — it ranks #9. The top model overall is GPT 5.5 (high) (OpenAI) at 22.2%.
Can open models match proprietary models on CL-bench Life?
Not quite on CL-bench Life: the strongest proprietary model (GPT 5.5 (high)) scores 22.2%, ahead of the best open model (DeepSeek V4 Pro) at 13.5% — but you can run the open one yourself.

Scores aggregated from epoch. llmrun does not run this benchmark — see the source for methodology, or the about benchmarks for what it measures.