Math

FrontierMath Tier 4 Leaderboard

FrontierMath Tier 4 is the hardest tier of Epoch AI's research-level mathematics benchmark, holding problems that working mathematicians consider genuinely difficult research questions. Scores here stay low even for frontier models, making it the sharpest available read on mathematical ability.

Source: epoch11 open models ranked+53 proprietaryData through Sep 2026

Open models ranked on FrontierMath Tier 4

# shows rank among open models / rank overall (including proprietary).

#ModelScore
1 / 24Kimi K3 · 2779.9B
39.0%
2 / 32GLM 5.2 · 753.3B
29.3%
3 / 33GLM 5.3 · 753.3B
29.3%
4 / 35DeepSeek V4 Pro 0813 · 1650.5B
26.8%
5 / 38Kimi K2.6 · 1026.9B
25.6%
6 / 39DeepSeek V4 Flash 0731 · 304.2B
24.4%
7 / 46GLM 5.3 Flash · 321.3B
17.1%
8 / 48Inkling Small · 266.0B
17.1%
9 / 52Kimi K2.7 Code · 1026.9B
12.2%
10 / 55Inkling · 952.4B
4.9%
11 / 59DeepSeek V4 Pro · 1598.8B
2.4%

Score vs model size

Which models give the most quality for their size — the ones worth running locally.

266B2.8Tmodel size (log scale) →39.0%2.4%GLM 5.3 · 753B · 29.3%DeepSeek V4 Pro 0813 · 1.7T · 26.8%Kimi K2.6 · 1T · 25.6%GLM 5.3 Flash · 321B · 17.1%Kimi K2.7 Code · 1T · 12.2%Inkling · 952B · 4.9%DeepSeek V4 Pro · 1.6T · 2.4%Inkling Small · 266B · 17.1%Inkling SmallDeepSeek V4 Flash 0731 · 304B · 24.4%DeepSeek V4 Flash 0731GLM 5.2 · 753B · 29.3%GLM 5.2Kimi K3 · 2.8T · 39.0%Kimi K3
Each dot is a model. Up = higher score, left = smaller (easier to run locally). The dashed line marks the efficiency frontier — the best score you can get at each size or smaller.
  • Inkling Small, 266B, score 17.1% — on the efficiency frontier (best score at its size or smaller).
  • DeepSeek V4 Flash 0731, 304B, score 24.4% — on the efficiency frontier (best score at its size or smaller).
  • GLM 5.2, 753B, score 29.3% — on the efficiency frontier (best score at its size or smaller).
  • Kimi K3, 2.8T, score 39.0% — on the efficiency frontier (best score at its size or smaller).

FrontierMath Tier 4: frequently asked questions

What is the best open LLM on FrontierMath Tier 4?
Kimi K3 is the top open model on FrontierMath Tier 4, scoring 39.0%. Among all models tested — including proprietary ones — it ranks #24. The top model overall is GPT 6 Astra (xhigh) (OpenAI) at 97.6%.
Can open models match proprietary models on FrontierMath Tier 4?
Not quite on FrontierMath Tier 4: the strongest proprietary model (GPT 6 Astra (xhigh)) scores 97.6%, ahead of the best open model (Kimi K3) at 39.0% — but you can run the open one yourself.

Scores aggregated from epoch. llmrun does not run this benchmark — see the source for methodology, or the about benchmarks for what it measures.