Reasoning

CL-bench Life Leaderboard

CL-bench Life, a follow-up from the CL-bench authors, tests the same context-learning skill on messier real-life material — fragmented notes, casual conversations, and behavioral logs — across 405 human-curated context-task pairs. It uses the same rubric-based LLM-judge scoring as CL-bench, requiring every rubric on a task to pass for it to count as solved.

Source: epoch6 open models ranked+11 proprietaryData through Apr 2026

All models ranked on CL-bench Life

Proprietary / closed models are shown dimmed — you can't run them locally, but they show where the open field stands.

#ModelScore
1GPT 5.5 (high) · proprietary
22.2%
2GPT 5.4 (Mar 05, 2026, xhigh) · proprietary
21.7%
3GPT 5.4 (Mar 05, 2026, high) · proprietary
19.3%
4GPT 5.1 (Nov 13, 2025, high) · proprietary
17.3%
5Claude Opus 4.6 (high) · proprietary
17.0%
6Gemini 3.1 Pro Preview · proprietary
16.9%
7GPT 5.4 (Mar 05, 2026, unspecified) · proprietary
13.8%
8Claude Opus 4.6 (unspecified) · proprietary
13.6%
9DeepSeek V4 Pro · 1598.8B
13.5%
10Kimi K2.5 · 1026.9B
13.2%
11Qwen3.5 Plus · proprietary
12.4%
12Grok 4.20 · proprietary
11.9%
13GLM 4.7 · 358.3B
10.9%
14DeepSeek V3.2 Exp · 685.4B
9.5%
15DeepSeek V3.2 · 685.4B
7.4%
16Mimo v2 Pro · proprietary
6.9%
17MiniMax M2.5 · 228.7B
6.3%

Score vs model size

Which models give the most quality for their size — the ones worth running locally.

229B1.6Tmodel size (log scale) →13.5%6.3%DeepSeek V3.2 Exp · 685B · 9.5%DeepSeek V3.2 · 685B · 7.4%MiniMax M2.5 · 229B · 6.3%MiniMax M2.5GLM 4.7 · 358B · 10.9%GLM 4.7Kimi K2.5 · 1T · 13.2%Kimi K2.5DeepSeek V4 Pro · 1.6T · 13.5%DeepSeek V4 Pro
Each dot is a model. Up = higher score, left = smaller (easier to run locally). The dashed line marks the efficiency frontier — the best score you can get at each size or smaller.
  • MiniMax M2.5, 229B, score 6.3% — on the efficiency frontier (best score at its size or smaller).
  • GLM 4.7, 358B, score 10.9% — on the efficiency frontier (best score at its size or smaller).
  • Kimi K2.5, 1T, score 13.2% — on the efficiency frontier (best score at its size or smaller).
  • DeepSeek V4 Pro, 1.6T, score 13.5% — on the efficiency frontier (best score at its size or smaller).

CL-bench Life: frequently asked questions

What is the best open LLM on CL-bench Life?
DeepSeek V4 Pro is the top open model on CL-bench Life, scoring 13.5%. Among all models tested — including proprietary ones — it ranks #9. The top model overall is GPT 5.5 (high) (OpenAI) at 22.2%.
Can open models match proprietary models on CL-bench Life?
Not quite on CL-bench Life: the strongest proprietary model (GPT 5.5 (high)) scores 22.2%, ahead of the best open model (DeepSeek V4 Pro) at 13.5% — but you can run the open one yourself.

Scores aggregated from epoch. llmrun does not run this benchmark — see the source for methodology, or the about benchmarks for what it measures.