CL-bench Life Leaderboard
CL-bench Life, a follow-up from the CL-bench authors, tests the same context-learning skill on messier real-life material — fragmented notes, casual conversations, and behavioral logs — across 405 human-curated context-task pairs. It uses the same rubric-based LLM-judge scoring as CL-bench, requiring every rubric on a task to pass for it to count as solved.
Source: epoch6 open models ranked+11 proprietaryData through Apr 2026
All models ranked on CL-bench Life
Proprietary / closed models are shown dimmed — you can't run them locally, but they show where the open field stands.
| # | Model | Score |
|---|---|---|
| 1 | GPT 5.5 (high) · proprietary | 22.2% |
| 2 | GPT 5.4 (Mar 05, 2026, xhigh) · proprietary | 21.7% |
| 3 | GPT 5.4 (Mar 05, 2026, high) · proprietary | 19.3% |
| 4 | GPT 5.1 (Nov 13, 2025, high) · proprietary | 17.3% |
| 5 | Claude Opus 4.6 (high) · proprietary | 17.0% |
| 6 | Gemini 3.1 Pro Preview · proprietary | 16.9% |
| 7 | GPT 5.4 (Mar 05, 2026, unspecified) · proprietary | 13.8% |
| 8 | Claude Opus 4.6 (unspecified) · proprietary | 13.6% |
| 9 | DeepSeek V4 Pro · 1598.8B | 13.5% |
| 10 | Kimi K2.5 · 1026.9B | 13.2% |
| 11 | Qwen3.5 Plus · proprietary | 12.4% |
| 12 | Grok 4.20 · proprietary | 11.9% |
| 13 | GLM 4.7 · 358.3B | 10.9% |
| 14 | DeepSeek V3.2 Exp · 685.4B | 9.5% |
| 15 | DeepSeek V3.2 · 685.4B | 7.4% |
| 16 | Mimo v2 Pro · proprietary | 6.9% |
| 17 | MiniMax M2.5 · 228.7B | 6.3% |
Score vs model size
Which models give the most quality for their size — the ones worth running locally.
- MiniMax M2.5, 229B, score 6.3% — on the efficiency frontier (best score at its size or smaller).
- GLM 4.7, 358B, score 10.9% — on the efficiency frontier (best score at its size or smaller).
- Kimi K2.5, 1T, score 13.2% — on the efficiency frontier (best score at its size or smaller).
- DeepSeek V4 Pro, 1.6T, score 13.5% — on the efficiency frontier (best score at its size or smaller).
CL-bench Life: frequently asked questions
- What is the best open LLM on CL-bench Life?
- DeepSeek V4 Pro is the top open model on CL-bench Life, scoring 13.5%. Among all models tested — including proprietary ones — it ranks #9. The top model overall is GPT 5.5 (high) (OpenAI) at 22.2%.
- Can open models match proprietary models on CL-bench Life?
- Not quite on CL-bench Life: the strongest proprietary model (GPT 5.5 (high)) scores 22.2%, ahead of the best open model (DeepSeek V4 Pro) at 13.5% — but you can run the open one yourself.
Scores aggregated from epoch. llmrun does not run this benchmark — see the source for methodology, or the about benchmarks for what it measures.