Knowledge
Humanity's Last Exam Leaderboard
Humanity's Last Exam (HLE) is a set of extremely difficult, expert-written questions across many fields, designed so that even frontier models score low. It is built to stay hard as models improve, measuring the true knowledge frontier.
Source: epoch4 open models ranked+43 proprietaryData through Sep 2026
All models ranked on Humanity's Last Exam
Proprietary / closed models are shown dimmed — you can't run them locally, but they show where the open field stands.
| # | Model | Score |
|---|---|---|
| 1 | Claude Fable 5.1 (xhigh) · proprietary | 46.5% |
| 2 | Gemini 3.1 Pro Preview · proprietary | 46.4% |
| 3 | GPT 5.4 Pro (Mar 05, 2026, unspecified) · proprietary | 44.3% |
| 4 | Muse Spark · proprietary | 40.6% |
| 5 | Gemini 3 Pro Preview · proprietary | 37.5% |
| 6 | GPT 5.4 (Mar 05, 2026, xhigh) · proprietary | 36.2% |
| 7 | Claude Opus 4.7 (unspecified) · proprietary | 36.2% |
| 8 | Claude Opus 4.6 Max · proprietary | 34.4% |
| 9 | GPT 5 Pro (Oct 06, 2025, unspecified) · proprietary | 31.6% |
| 10 | GPT 5.2 (Dec 11, 2025, unspecified) · proprietary | 27.8% |
| 11 | GPT 5 (Aug 07, 2025, high) · proprietary | 25.3% |
| 12 | GPT 5 (Aug 07, 2025, unspecified) · proprietary | 25.3% |
| 13 | Claude Opus 4.5 (Nov 01, 2025, unspecified) · proprietary | 25.2% |
| 14 | Kimi K2.5 · 1026.9B | 24.4% |
| 15 | GPT 5.1 (Nov 13, 2025, unspecified) · proprietary | 23.7% |
| 16 | Gemini 2.5 Pro Preview (Jun 05) · proprietary | 21.6% |
| 17 | O3 (Apr 16, 2025, high) · proprietary | 20.3% |
| 18 | GPT 5 Mini (Aug 07, 2025, unspecified) · proprietary | 19.4% |
| 19 | O3 (Apr 16, 2025, medium) · proprietary | 19.2% |
| 20 | Claude Opus 4.6 · proprietary | 19.0% |
| 21 | Gemini 2.5 Pro Exp (Mar 25) · proprietary | 18.2% |
| 22 | O4 Mini (Apr 16, 2025, high) · proprietary | 18.1% |
| 23 | Gemini 2.5 Pro Preview (May 06) · proprietary | 17.8% |
| 24 | O4 Mini (Apr 16, 2025, medium) · proprietary | 14.3% |
| 25 | Claude Sonnet 4.5 (Sep 29, 2025, unspecified) · proprietary | 13.7% |
| 26 | Gemini 2.5 Flash Preview (Apr 17) · proprietary | 12.1% |
| 27 | Claude Opus 4.1 (Aug 05, 2025, unspecified) · proprietary | 11.5% |
| 28 | Gemini 2.5 Flash Preview (May 20) · proprietary | 11.0% |
| 29 | Claude Opus 4 (May 14, 2025, unspecified) · proprietary | 10.7% |
| 30 | Gemini 3.1 Flash Lite · proprietary | 8.6% |
| 31 | GLM 4.5 · 358.3B | 8.3% |
| 32 | GLM 4.5 Air · 110.5B | 8.1% |
| 33 | O1 Pro (Mar 19, 2025) · proprietary | 8.1% |
| 34 | Claude 3.7 Sonnet (Feb 19, 2025, unspecified) · proprietary | 8.0% |
| 35 | O1 (Dec 17, 2024, unspecified) · proprietary | 8.0% |
| 36 | Claude Sonnet 4 (May 14, 2025, unspecified) · proprietary | 7.8% |
| 37 | GPT 5.1 2025 11.13 None · proprietary | 6.8% |
| 38 | Gemini 2.0 Flash Thinking Exp (Jan 21) · proprietary | 6.6% |
| 39 | Llama 4 Maverick 17B 128E Instruct · 401.6B | 5.7% |
| 40 | GPT 4.5 Preview (Feb 27, 2025) · proprietary | 5.4% |
| 41 | GPT 4.1 (Apr 14, 2025) · proprietary | 5.4% |
| 42 | Gemini 1.5 Pro 002 · proprietary | 4.6% |
| 43 | Mistral Medium 2505 · proprietary | 4.5% |
| 44 | Amazon.nova Pro v1:0 · proprietary | 4.4% |
| 45 | Claude 3.5 Sonnet (Oct 22, 2024) · proprietary | 4.1% |
| 46 | Amazon.nova Lite v1:0 · proprietary | 3.6% |
| 47 | GPT 4o (Nov 20, 2024) · proprietary | 2.7% |
Humanity's Last Exam: frequently asked questions
- What is the best open LLM on Humanity's Last Exam?
- Kimi K2.5 is the top open model on Humanity's Last Exam, scoring 24.4%. Among all models tested — including proprietary ones — it ranks #14. The top model overall is Claude Fable 5.1 (xhigh) (Anthropic) at 46.5%.
- Can open models match proprietary models on Humanity's Last Exam?
- Not quite on Humanity's Last Exam: the strongest proprietary model (Claude Fable 5.1 (xhigh)) scores 46.5%, ahead of the best open model (Kimi K2.5) at 24.4% — but you can run the open one yourself.
Scores aggregated from epoch. llmrun does not run this benchmark — see the source for methodology, or the about benchmarks for what it measures.