Math
FrontierMath Tier 4 Leaderboard
FrontierMath Tier 4 is the hardest tier of Epoch AI's research-level mathematics benchmark, holding problems that working mathematicians consider genuinely difficult research questions. Scores here stay low even for frontier models, making it the sharpest available read on mathematical ability.
Source: epoch11 open models ranked+53 proprietaryData through Sep 2026
All models ranked on FrontierMath Tier 4
Proprietary / closed models are shown dimmed — you can't run them locally, but they show where the open field stands.
| # | Model | Score |
|---|---|---|
| 1 | GPT 6 Astra (high) · proprietary | 97.6% |
| 2 | GPT 6 Astra (xhigh) · proprietary | 97.6% |
| 3 | GPT 6 Astra Max · proprietary | 97.6% |
| 4 | GPT 6 Astra (medium) · proprietary | 97.6% |
| 5 | Claude Fable 5 Max · proprietary | 90.2% |
| 6 | Claude Fable 5.1 Max · proprietary | 87.8% |
| 7 | GPT 6 Astra (low) · proprietary | 87.8% |
| 8 | GPT 5.6 Sol Max · proprietary | 82.9% |
| 9 | GPT 6 Astra None · proprietary | 82.9% |
| 10 | GPT 5.6 Sol Promax · proprietary | 80.5% |
| 11 | GPT 5.5 Pro (xhigh) · proprietary | 78.0% |
| 12 | Gdm Ai Co Mathematician · proprietary | 75.6% |
| 13 | Claude Opus 5 Max · proprietary | 73.2% |
| 14 | GPT 5.5 (xhigh) · proprietary | 72.5% |
| 15 | GPT 5.6 Terra Max · proprietary | 70.7% |
| 16 | GPT 5.6 Luna Max · proprietary | 61.0% |
| 17 | GPT 5.4 Pro (Mar 05, 2026, xhigh) · proprietary | 58.5% |
| 18 | Claude Opus 4.8 Max · proprietary | 56.1% |
| 19 | GPT 5.4 (Mar 05, 2026, xhigh) · proprietary | 49.0% |
| 20 | Muse Spark 1.3 Max · proprietary | 46.3% |
| 21 | Qwen3.8 Max (xhigh) · proprietary | 46.3% |
| 22 | GPT 5.2 Pro (Dec 11, 2025, xhigh) · proprietary | 46.0% |
| 23 | Muse Spark 1.3 (xhigh) · proprietary | 41.5% |
| 24 | Kimi K3 · 2779.9B | 39.0% |
| 25 | Gemini 3.7 Flash (high) · proprietary | 36.6% |
| 26 | Qwen3.7 Max · proprietary | 34.2% |
| 27 | Qwen3.8 Max (Sep 02, xhigh) · proprietary | 34.2% |
| 28 | Claude Opus 4.7 Max · proprietary | 31.7% |
| 29 | Grok 4.6 (xhigh) · proprietary | 31.7% |
| 30 | GPT 5.2 (Dec 11, 2025, xhigh) · proprietary | 31.7% |
| 31 | Claude Sonnet 5 Max · proprietary | 29.3% |
| 32 | GLM 5.2 · 753.3B | 29.3% |
| 33 | GLM 5.3 · 753.3B | 29.3% |
| 34 | Claude Opus 4.6 Max · proprietary | 26.8% |
| 35 | DeepSeek V4 Pro 0813 · 1650.5B | 26.8% |
| 36 | Gemini 3.1 Pro Preview · proprietary | 26.8% |
| 37 | Gemini 3.5 Flash (high) · proprietary | 26.8% |
| 38 | Kimi K2.6 · 1026.9B | 25.6% |
| 39 | DeepSeek V4 Flash 0731 · 304.2B | 24.4% |
| 40 | Grok 4.5 (high) · proprietary | 24.4% |
| 41 | Gemini 3.6 Flash (high) · proprietary | 21.9% |
| 42 | Gemini 3.8 Flash (high) · proprietary | 21.9% |
| 43 | GPT 5 (Aug 07, 2025, high) · proprietary | 21.9% |
| 44 | GPT 5 Pro (Oct 06, 2025, high) · proprietary | 19.5% |
| 45 | Gemini 3 Flash Preview · proprietary | 17.1% |
| 46 | GLM 5.3 Flash · 321.3B | 17.1% |
| 47 | Grok 4.20 0309 Reasoning · proprietary | 17.1% |
| 48 | Inkling Small · 266.0B | 17.1% |
| 49 | Grok 4.3 (high) · proprietary | 14.6% |
| 50 | GPT 5 Mini (Aug 07, 2025, high) · proprietary | 12.2% |
| 51 | GPT 5.4 Nano (Mar 17, 2026, high) · proprietary | 12.2% |
| 52 | Kimi K2.7 Code · 1026.9B | 12.2% |
| 53 | GPT 5.4 Mini (Mar 17, 2026, xhigh) · proprietary | 9.8% |
| 54 | Claude Opus 4.5 (Nov 01, 2025, 32K) · proprietary | 4.9% |
| 55 | Inkling · 952.4B | 4.9% |
| 56 | O4 Mini (Apr 16, 2025, high) · proprietary | 4.9% |
| 57 | Claude Opus 4.1 (Aug 05, 2025, 32K) · proprietary | 2.4% |
| 58 | Claude Sonnet 4.5 (Sep 29, 2025, 32K) · proprietary | 2.4% |
| 59 | DeepSeek V4 Pro · 1598.8B | 2.4% |
| 60 | GPT 5 Nano (Aug 07, 2025, high) · proprietary | 2.4% |
| 61 | GPT 5.5 Instant · proprietary | 2.4% |
| 62 | Gemini 2.5 Pro · proprietary | 0.0% |
| 63 | Gemini 3.5 Flash Lite (high) · proprietary | 0.0% |
| 64 | O3 Mini (Jan 31, 2025, high) · proprietary | 0.0% |
Score vs model size
Which models give the most quality for their size — the ones worth running locally.
- Inkling Small, 266B, score 17.1% — on the efficiency frontier (best score at its size or smaller).
- DeepSeek V4 Flash 0731, 304B, score 24.4% — on the efficiency frontier (best score at its size or smaller).
- GLM 5.2, 753B, score 29.3% — on the efficiency frontier (best score at its size or smaller).
- Kimi K3, 2.8T, score 39.0% — on the efficiency frontier (best score at its size or smaller).
FrontierMath Tier 4: frequently asked questions
- What is the best open LLM on FrontierMath Tier 4?
- Kimi K3 is the top open model on FrontierMath Tier 4, scoring 39.0%. Among all models tested — including proprietary ones — it ranks #24. The top model overall is GPT 6 Astra (xhigh) (OpenAI) at 97.6%.
- Can open models match proprietary models on FrontierMath Tier 4?
- Not quite on FrontierMath Tier 4: the strongest proprietary model (GPT 6 Astra (xhigh)) scores 97.6%, ahead of the best open model (Kimi K3) at 39.0% — but you can run the open one yourself.
Scores aggregated from epoch. llmrun does not run this benchmark — see the source for methodology, or the about benchmarks for what it measures.