Math

FrontierMath Tier 4 Leaderboard

FrontierMath Tier 4 is the hardest tier of Epoch AI's research-level mathematics benchmark, holding problems that working mathematicians consider genuinely difficult research questions. Scores here stay low even for frontier models, making it the sharpest available read on mathematical ability.

Source: epoch11 open models ranked+53 proprietaryData through Sep 2026

All models ranked on FrontierMath Tier 4

Proprietary / closed models are shown dimmed — you can't run them locally, but they show where the open field stands.

#ModelScore
1GPT 6 Astra (high) · proprietary
97.6%
2GPT 6 Astra (xhigh) · proprietary
97.6%
3GPT 6 Astra Max · proprietary
97.6%
4GPT 6 Astra (medium) · proprietary
97.6%
5Claude Fable 5 Max · proprietary
90.2%
6Claude Fable 5.1 Max · proprietary
87.8%
7GPT 6 Astra (low) · proprietary
87.8%
8GPT 5.6 Sol Max · proprietary
82.9%
9GPT 6 Astra None · proprietary
82.9%
10GPT 5.6 Sol Promax · proprietary
80.5%
11GPT 5.5 Pro (xhigh) · proprietary
78.0%
12Gdm Ai Co Mathematician · proprietary
75.6%
13Claude Opus 5 Max · proprietary
73.2%
14GPT 5.5 (xhigh) · proprietary
72.5%
15GPT 5.6 Terra Max · proprietary
70.7%
16GPT 5.6 Luna Max · proprietary
61.0%
17GPT 5.4 Pro (Mar 05, 2026, xhigh) · proprietary
58.5%
18Claude Opus 4.8 Max · proprietary
56.1%
19GPT 5.4 (Mar 05, 2026, xhigh) · proprietary
49.0%
20Muse Spark 1.3 Max · proprietary
46.3%
21Qwen3.8 Max (xhigh) · proprietary
46.3%
22GPT 5.2 Pro (Dec 11, 2025, xhigh) · proprietary
46.0%
23Muse Spark 1.3 (xhigh) · proprietary
41.5%
24Kimi K3 · 2779.9B
39.0%
25Gemini 3.7 Flash (high) · proprietary
36.6%
26Qwen3.7 Max · proprietary
34.2%
27Qwen3.8 Max (Sep 02, xhigh) · proprietary
34.2%
28Claude Opus 4.7 Max · proprietary
31.7%
29Grok 4.6 (xhigh) · proprietary
31.7%
30GPT 5.2 (Dec 11, 2025, xhigh) · proprietary
31.7%
31Claude Sonnet 5 Max · proprietary
29.3%
32GLM 5.2 · 753.3B
29.3%
33GLM 5.3 · 753.3B
29.3%
34Claude Opus 4.6 Max · proprietary
26.8%
35DeepSeek V4 Pro 0813 · 1650.5B
26.8%
36Gemini 3.1 Pro Preview · proprietary
26.8%
37Gemini 3.5 Flash (high) · proprietary
26.8%
38Kimi K2.6 · 1026.9B
25.6%
39DeepSeek V4 Flash 0731 · 304.2B
24.4%
40Grok 4.5 (high) · proprietary
24.4%
41Gemini 3.6 Flash (high) · proprietary
21.9%
42Gemini 3.8 Flash (high) · proprietary
21.9%
43GPT 5 (Aug 07, 2025, high) · proprietary
21.9%
44GPT 5 Pro (Oct 06, 2025, high) · proprietary
19.5%
45Gemini 3 Flash Preview · proprietary
17.1%
46GLM 5.3 Flash · 321.3B
17.1%
47Grok 4.20 0309 Reasoning · proprietary
17.1%
48Inkling Small · 266.0B
17.1%
49Grok 4.3 (high) · proprietary
14.6%
50GPT 5 Mini (Aug 07, 2025, high) · proprietary
12.2%
51GPT 5.4 Nano (Mar 17, 2026, high) · proprietary
12.2%
52Kimi K2.7 Code · 1026.9B
12.2%
53GPT 5.4 Mini (Mar 17, 2026, xhigh) · proprietary
9.8%
54Claude Opus 4.5 (Nov 01, 2025, 32K) · proprietary
4.9%
55Inkling · 952.4B
4.9%
56O4 Mini (Apr 16, 2025, high) · proprietary
4.9%
57Claude Opus 4.1 (Aug 05, 2025, 32K) · proprietary
2.4%
58Claude Sonnet 4.5 (Sep 29, 2025, 32K) · proprietary
2.4%
59DeepSeek V4 Pro · 1598.8B
2.4%
60GPT 5 Nano (Aug 07, 2025, high) · proprietary
2.4%
61GPT 5.5 Instant · proprietary
2.4%
62Gemini 2.5 Pro · proprietary
0.0%
63Gemini 3.5 Flash Lite (high) · proprietary
0.0%
64O3 Mini (Jan 31, 2025, high) · proprietary
0.0%

Score vs model size

Which models give the most quality for their size — the ones worth running locally.

266B2.8Tmodel size (log scale) →39.0%2.4%GLM 5.3 · 753B · 29.3%DeepSeek V4 Pro 0813 · 1.7T · 26.8%Kimi K2.6 · 1T · 25.6%GLM 5.3 Flash · 321B · 17.1%Kimi K2.7 Code · 1T · 12.2%Inkling · 952B · 4.9%DeepSeek V4 Pro · 1.6T · 2.4%Inkling Small · 266B · 17.1%Inkling SmallDeepSeek V4 Flash 0731 · 304B · 24.4%DeepSeek V4 Flash 0731GLM 5.2 · 753B · 29.3%GLM 5.2Kimi K3 · 2.8T · 39.0%Kimi K3
Each dot is a model. Up = higher score, left = smaller (easier to run locally). The dashed line marks the efficiency frontier — the best score you can get at each size or smaller.
  • Inkling Small, 266B, score 17.1% — on the efficiency frontier (best score at its size or smaller).
  • DeepSeek V4 Flash 0731, 304B, score 24.4% — on the efficiency frontier (best score at its size or smaller).
  • GLM 5.2, 753B, score 29.3% — on the efficiency frontier (best score at its size or smaller).
  • Kimi K3, 2.8T, score 39.0% — on the efficiency frontier (best score at its size or smaller).

FrontierMath Tier 4: frequently asked questions

What is the best open LLM on FrontierMath Tier 4?
Kimi K3 is the top open model on FrontierMath Tier 4, scoring 39.0%. Among all models tested — including proprietary ones — it ranks #24. The top model overall is GPT 6 Astra (xhigh) (OpenAI) at 97.6%.
Can open models match proprietary models on FrontierMath Tier 4?
Not quite on FrontierMath Tier 4: the strongest proprietary model (GPT 6 Astra (xhigh)) scores 97.6%, ahead of the best open model (Kimi K3) at 39.0% — but you can run the open one yourself.

Scores aggregated from epoch. llmrun does not run this benchmark — see the source for methodology, or the about benchmarks for what it measures.

FrontierMath Tier 4 Leaderboard — LLM Scores | llmrun