Math

FrontierMath Leaderboard

FrontierMath is a benchmark of exceptionally hard, original research-level mathematics problems created with professional mathematicians. Even the strongest models solve only a small fraction, making it a frontier measure of genuine mathematical ability.

Source: epoch12 open models ranked+89 proprietaryData through May 2026

Open models ranked on FrontierMath

# shows rank among open models / rank overall (including proprietary).

#ModelScore
1 / 15Kimi K2.6 · 1058.6B
39.0%
2 / 21GLM 5.1 · 753.9B
33.5%
3 / 27Kimi K2.5 · 1058.6B
27.9%
4 / 36DeepSeek V3.2 · 685.4B
22.1%
5 / 37Kimi K2 Thinking · 1058.1B
21.4%
6 / 48GLM 5 · 753.9B
16.4%
7 / 59Qwen3 235B A22B Thinking 2507 · 235.1B
8.5%
8 / 76GLM 4.6 · 356.8B
3.8%
9 / 82GLM 4.7 · 358.3B
2.4%
10 / 85DeepSeek v3 · 684.5B
1.7%
11 / 94Llama 4 Maverick 17B 128E Instruct · 401.6B
0.7%
12 / 101Llama 4 Scout 17B 16E Instruct · 108.6B
0.0%

Score vs model size

Which models give the most quality for their size — the ones worth running locally.

109B1.1Tmodel size (log scale) →39.0%0.0%Kimi K2.5 · 1.1T · 27.9%Kimi K2 Thinking · 1.1T · 21.4%GLM 5 · 754B · 16.4%GLM 4.6 · 357B · 3.8%GLM 4.7 · 358B · 2.4%DeepSeek v3 · 685B · 1.7%Llama 4 Maverick 17B 128E Instruct · 402B · 0.7%Llama 4 Scout 17B 16E Instruct · 109B · 0.0%Llama 4 Scout 17B 16E…Qwen3 235B A22B Thinking 2507 · 235B · 8.5%Qwen3 235B A22B Think…DeepSeek V3.2 · 685B · 22.1%DeepSeek V3.2GLM 5.1 · 754B · 33.5%GLM 5.1Kimi K2.6 · 1.1T · 39.0%Kimi K2.6
Each dot is a model. Up = higher score, left = smaller (easier to run locally). The dashed line marks the efficiency frontier — the best score you can get at each size or smaller.
  • Llama 4 Scout 17B 16E Instruct, 109B, score 0.0% — on the efficiency frontier (best score at its size or smaller).
  • Qwen3 235B A22B Thinking 2507, 235B, score 8.5% — on the efficiency frontier (best score at its size or smaller).
  • DeepSeek V3.2, 685B, score 22.1% — on the efficiency frontier (best score at its size or smaller).
  • GLM 5.1, 754B, score 33.5% — on the efficiency frontier (best score at its size or smaller).
  • Kimi K2.6, 1.1T, score 39.0% — on the efficiency frontier (best score at its size or smaller).

FrontierMath: frequently asked questions

What is the best open LLM on FrontierMath?
Kimi K2.6 is the top open model on FrontierMath, scoring 39.0%. Among all models tested — including proprietary ones — it ranks #14. The top model overall is GPT 5.5 Pro Pre Release (high) (OpenAI) at 52.4%.
Can open models match proprietary models on FrontierMath?
Not quite on FrontierMath: the strongest proprietary model (GPT 5.5 Pro Pre Release (high)) scores 52.4%, ahead of the best open model (Kimi K2.6) at 39.0% — but you can run the open one yourself.

Scores aggregated from epoch. llmrun does not run this benchmark — see the source for methodology, or the about benchmarks for what it measures.