Math
FrontierMath Tier 4 Leaderboard
FrontierMath Tier 4 is the hardest tier of Epoch AI's research-level mathematics benchmark, holding problems that working mathematicians consider genuinely difficult research questions. Scores here stay low even for frontier models, making it the sharpest available read on mathematical ability.
Source: epoch11 open models ranked+53 proprietaryData through Sep 2026
Open models ranked on FrontierMath Tier 4
# shows rank among open models / rank overall (including proprietary).
| # | Model | Score |
|---|---|---|
| 1 / 24 | Kimi K3 · 2779.9B | 39.0% |
| 2 / 32 | GLM 5.2 · 753.3B | 29.3% |
| 3 / 33 | GLM 5.3 · 753.3B | 29.3% |
| 4 / 35 | DeepSeek V4 Pro 0813 · 1650.5B | 26.8% |
| 5 / 38 | Kimi K2.6 · 1026.9B | 25.6% |
| 6 / 39 | DeepSeek V4 Flash 0731 · 304.2B | 24.4% |
| 7 / 46 | GLM 5.3 Flash · 321.3B | 17.1% |
| 8 / 48 | Inkling Small · 266.0B | 17.1% |
| 9 / 52 | Kimi K2.7 Code · 1026.9B | 12.2% |
| 10 / 55 | Inkling · 952.4B | 4.9% |
| 11 / 59 | DeepSeek V4 Pro · 1598.8B | 2.4% |
Score vs model size
Which models give the most quality for their size — the ones worth running locally.
- Inkling Small, 266B, score 17.1% — on the efficiency frontier (best score at its size or smaller).
- DeepSeek V4 Flash 0731, 304B, score 24.4% — on the efficiency frontier (best score at its size or smaller).
- GLM 5.2, 753B, score 29.3% — on the efficiency frontier (best score at its size or smaller).
- Kimi K3, 2.8T, score 39.0% — on the efficiency frontier (best score at its size or smaller).
FrontierMath Tier 4: frequently asked questions
- What is the best open LLM on FrontierMath Tier 4?
- Kimi K3 is the top open model on FrontierMath Tier 4, scoring 39.0%. Among all models tested — including proprietary ones — it ranks #24. The top model overall is GPT 6 Astra (xhigh) (OpenAI) at 97.6%.
- Can open models match proprietary models on FrontierMath Tier 4?
- Not quite on FrontierMath Tier 4: the strongest proprietary model (GPT 6 Astra (xhigh)) scores 97.6%, ahead of the best open model (Kimi K3) at 39.0% — but you can run the open one yourself.
Scores aggregated from epoch. llmrun does not run this benchmark — see the source for methodology, or the about benchmarks for what it measures.