Knowledge

Humanity's Last Exam Leaderboard

Humanity's Last Exam (HLE) is a set of extremely difficult, expert-written questions across many fields, designed so that even frontier models score low. It is built to stay hard as models improve, measuring the true knowledge frontier.

Source: epoch4 open models ranked+42 proprietaryData through Apr 2026

All models ranked on Humanity's Last Exam

Proprietary / closed models are shown dimmed — you can't run them locally, but they show where the open field stands.

#ModelScore
1Gemini 3.1 Pro Preview · proprietary
46.4%
2GPT 5.4 Pro (Mar 05, 2026, unspecified) · proprietary
44.3%
3Muse Spark · proprietary
40.6%
4Gemini 3 Pro Preview · proprietary
37.5%
5GPT 5.4 (Mar 05, 2026, xhigh) · proprietary
36.2%
6Claude Opus 4.7 (unspecified) · proprietary
36.2%
7Claude Opus 4.6 Max · proprietary
34.4%
8GPT 5 Pro (Oct 06, 2025, unspecified) · proprietary
31.6%
9GPT 5.2 (Dec 11, 2025, unspecified) · proprietary
27.8%
10GPT 5 (Aug 07, 2025, high) · proprietary
25.3%
11GPT 5 (Aug 07, 2025, unspecified) · proprietary
25.3%
12Claude Opus 4.5 (Nov 01, 2025, unspecified) · proprietary
25.2%
13Kimi K2.5 · 1058.6B
24.4%
14GPT 5.1 (Nov 13, 2025, unspecified) · proprietary
23.7%
15Gemini 2.5 Pro Preview (Jun 05) · proprietary
21.6%
16O3 (Apr 16, 2025, high) · proprietary
20.3%
17GPT 5 Mini (Aug 07, 2025, unspecified) · proprietary
19.4%
18O3 (Apr 16, 2025, medium) · proprietary
19.2%
19Claude Opus 4.6 · proprietary
19.0%
20Gemini 2.5 Pro Exp (Mar 25) · proprietary
18.2%
21O4 Mini (Apr 16, 2025, high) · proprietary
18.1%
22Gemini 2.5 Pro Preview (May 06) · proprietary
17.8%
23O4 Mini (Apr 16, 2025, medium) · proprietary
14.3%
24Claude Sonnet 4.5 (Sep 29, 2025, unspecified) · proprietary
13.7%
25Gemini 2.5 Flash Preview (Apr 17) · proprietary
12.1%
26Claude Opus 4.1 (Aug 05, 2025, unspecified) · proprietary
11.5%
27Gemini 2.5 Flash Preview (May 20) · proprietary
11.0%
28Claude Opus 4 (May 14, 2025, unspecified) · proprietary
10.7%
29Gemini 3.1 Flash Lite · proprietary
8.6%
30GLM 4.5 · 358.3B
8.3%
31GLM 4.5 Air · 110.5B
8.1%
32O1 Pro (Mar 19, 2025) · proprietary
8.1%
33Claude 3.7 Sonnet (Feb 19, 2025, unspecified) · proprietary
8.0%
34O1 (Dec 17, 2024, unspecified) · proprietary
8.0%
35Claude Sonnet 4 (May 14, 2025, unspecified) · proprietary
7.8%
36GPT 5.1 2025 11.13 None · proprietary
6.8%
37Gemini 2.0 Flash Thinking Exp (Jan 21) · proprietary
6.6%
38Llama 4 Maverick 17B 128E Instruct · 401.6B
5.7%
39GPT 4.5 Preview (Feb 27, 2025) · proprietary
5.4%
40GPT 4.1 (Apr 14, 2025) · proprietary
5.4%
41Gemini 1.5 Pro 002 · proprietary
4.6%
42Mistral Medium 2505 · proprietary
4.5%
43Amazon.nova Pro v1:0 · proprietary
4.4%
44Claude 3.5 Sonnet (Oct 22, 2024) · proprietary
4.1%
45Amazon.nova Lite v1:0 · proprietary
3.6%
46GPT 4o (Nov 20, 2024) · proprietary
2.7%

Humanity's Last Exam: frequently asked questions

What is the best open LLM on Humanity's Last Exam?
Kimi K2.5 is the top open model on Humanity's Last Exam, scoring 24.4%. Among all models tested — including proprietary ones — it ranks #13. The top model overall is Gemini 3.1 Pro Preview (Google DeepMind) at 46.4%.
Can open models match proprietary models on Humanity's Last Exam?
Not quite on Humanity's Last Exam: the strongest proprietary model (Gemini 3.1 Pro Preview) scores 46.4%, ahead of the best open model (Kimi K2.5) at 24.4% — but you can run the open one yourself.

Scores aggregated from epoch. llmrun does not run this benchmark — see the source for methodology, or the about benchmarks for what it measures.