Math

AIME 2024/2025 Leaderboard

AIME (American Invitational Mathematics Examination) is a prestigious high-school competition of hard, integer-answer problems. It's a widely-cited yardstick for multi-step mathematical reasoning, here on the 2024–2025 papers.

Source: epoch74 open models ranked+196 proprietaryData through Sep 2026

Open models ranked on AIME 2024/2025

# shows rank among open models / rank overall (including proprietary).

#ModelScore
1 / 12DeepSeek V4 Pro 0813 · 1650.5B
98.6%
2 / 23Kimi K3 · 2779.9B
97.2%
3 / 24DeepSeek V4 Pro · 1598.8B
96.7%
4 / 26Kimi K2.6 · 1026.9B
96.1%
5 / 34Kimi K2.7 Code · 1026.9B
95.6%
6 / 39DeepSeek V4 Flash 0731 · 304.2B
94.4%
7 / 41GLM 5.3 Flash · 321.3B
93.9%
8 / 44GLM 5.1 · 753.9B
93.3%
9 / 51Kimi K2.5 · 1026.9B
92.2%
10 / 55GLM 5.3 · 753.3B
91.1%
11 / 56Qwen3.6 27B · 27.8B
91.1%
12 / 58Inkling Small · 266.0B
90.0%
13 / 63GPT OSS 120B · 116.8B
88.9%
14 / 64Inkling · 952.4B
88.9%
15 / 65Qwen3.5 397B A17B · 403.4B
88.9%
16 / 74Qwen3 235B A22B Thinking 2507 · 235.1B
86.7%
17 / 76Qwen3.6 35B A3B · 36.0B
86.7%
18 / 78GLM 5.2 · 753.3B
86.4%
19 / 91GLM 4.7 · 358.3B
83.3%
20 / 92Kimi K2 Thinking · 1026.4B
83.1%
21 / 95Gemma 4 26B A4B IT · 25.8B
82.2%
22 / 103GLM 5 · 753.9B
80.0%
23 / 114Gemma 4 31B IT · 31.3B
73.3%
24 / 124MiniMax M3 · 427.0B
71.1%
25 / 126Qwen3 30B A3B Thinking 2507 · 30.5B
70.3%
26 / 133Seed OSS 36B Instruct · 36.2B
67.5%
27 / 134Qwen3 32B · 32.8B
66.9%
28 / 137DeepSeek R1 0528 · 684.5B
66.4%
29 / 138Qwen3 14B · 14.8B
66.4%
30 / 139GPT OSS 20B · 20.9B
65.3%
31 / 144Qwen3 30B A3B · 30.5B
62.8%
32 / 147Qwen3 30B A3B Instruct 2507 · 30.5B
62.2%
33 / 148Qwen3.5 9B · 9.7B
61.7%
34 / 152QwQ 32B · 32.8B
59.2%
35 / 153GLM 4.7 Flash · 31.2B
58.3%
36 / 159Qwen3 8B · 8.2B
56.1%
37 / 160Qwen3.5 4B · 4.7B
55.8%
38 / 161DeepSeek R1 Distill Qwen 32B · 32.8B
55.6%
39 / 167DeepSeek R1 · 684.5B
53.3%
40 / 170Qwen3 4B Instruct 2507 · 4.0B
52.2%
41 / 171DeepSeek R1 Distill Llama 70B · 70.6B
51.4%
42 / 173DeepSeek R1 Distill Qwen 14B · 14.8B
50.6%
43 / 185DeepSeek R1 0528 Qwen3 8B · 8.2B
43.9%
44 / 190DeepSeek v3 0324 · 684.5B
37.8%
45 / 201Magistral Small 2506 · 23.6B
30.0%
46 / 207Gemma 3 27B IT · 27.4B
22.5%
47 / 209DeepSeek R1 Distill Qwen 1.5B · 1.8B
21.4%
48 / 210Llama 4 Maverick 17B 128E Instruct · 401.6B
20.6%
49 / 212Gemma 3 12B IT · 12.2B
16.7%
50 / 215DeepSeek v3 · 684.5B
15.8%
51 / 216Phi 4 · 14.7B
13.8%
52 / 218Llama 3.1 405B Instruct · 405.9B
9.7%
53 / 221Qwen2.5 72B Instruct · 72.7B
8.1%
54 / 222Qwen3 1.7B · 2.0B
8.1%
55 / 223Llama 4 Scout 17B 16E Instruct · 108.6B
7.8%
56 / 225Gemma 3 4B IT · 4.3B
7.5%
57 / 226Qwen2.5 32B Instruct · 32.8B
7.4%
58 / 238Llama 3.3 70B Instruct · 70.6B
5.1%
59 / 241Llama 3.1 Tulu 3 70B DPO · 70.6B
4.4%
60 / 243Meta Llama 3 70B Instruct · 70.6B
4.3%
61 / 246Llama 3.1 70B Instruct · 70.6B
3.6%
62 / 247Granite 4.0 Micro · 3.4B
2.8%
63 / 248Llama 3.2 90B Vision Instruct · 88.6B
2.6%
64 / 251Hermes 2 Theta Llama 3 70B · 70.6B
2.5%
65 / 252Qwen2.5 7B Instruct · 7.6B
2.5%
66 / 255Meta Llama 3 8B Instruct · 8.0B
1.9%
67 / 258Llama 3.1 8B Instruct · 8.0B
1.7%
68 / 259Gemma 2 27B IT · 27.2B
1.4%
69 / 261Gemma 3 1B IT · 1000M
1.1%
70 / 263Deepseek Llm 67B Chat · 67B
0.8%
71 / 265Gemma 2 9B IT · 9.2B
0.6%
72 / 267Llama 3.2 1B Instruct · 1.2B
0.6%
73 / 269Mistral 7B Instruct v0.3 · 7.2B
0.3%
74 / 270Llama 2 70B Chat HF · 69.0B
0.0%

Score vs model size

Which models give the most quality for their size — the ones worth running locally.

1B10B100B1Tmodel size (log scale) →98.6%0.0%Kimi K3 · 2.8T · 97.2%Kimi K2.7 Code · 1T · 95.6%GLM 5.3 Flash · 321B · 93.9%GLM 5.1 · 754B · 93.3%Kimi K2.5 · 1T · 92.2%GLM 5.3 · 753B · 91.1%Inkling Small · 266B · 90.0%GPT OSS 120B · 117B · 88.9%Qwen3.5 397B A17B · 403B · 88.9%Inkling · 952B · 88.9%Qwen3 235B A22B Thinking 2507 · 235B · 86.7%Qwen3.6 35B A3B · 36B · 86.7%GLM 5.2 · 753B · 86.4%GLM 4.7 · 358B · 83.3%Kimi K2 Thinking · 1T · 83.1%GLM 5 · 754B · 80.0%Gemma 4 31B IT · 31B · 73.3%MiniMax M3 · 427B · 71.1%Qwen3 30B A3B Thinking 2507 · 31B · 70.3%Seed OSS 36B Instruct · 36B · 67.5%Qwen3 32B · 33B · 66.9%DeepSeek R1 0528 · 685B · 66.4%GPT OSS 20B · 21B · 65.3%Qwen3 30B A3B · 31B · 62.8%Qwen3 30B A3B Instruct 2507 · 31B · 62.2%QwQ 32B · 33B · 59.2%GLM 4.7 Flash · 31B · 58.3%DeepSeek R1 Distill Qwen 32B · 33B · 55.6%DeepSeek R1 · 685B · 53.3%DeepSeek R1 Distill Llama 70B · 71B · 51.4%DeepSeek R1 Distill Qwen 14B · 15B · 50.6%DeepSeek R1 0528 Qwen3 8B · 8B · 43.9%DeepSeek v3 0324 · 685B · 37.8%Magistral Small 2506 · 24B · 30.0%Gemma 3 27B IT · 27B · 22.5%Llama 4 Maverick 17B 128E Instruct · 402B · 20.6%Gemma 3 12B IT · 12B · 16.7%DeepSeek v3 · 685B · 15.8%Phi 4 · 15B · 13.8%Llama 3.1 405B Instruct · 406B · 9.7%Qwen2.5 72B Instruct · 73B · 8.1%Qwen3 1.7B · 2B · 8.1%Llama 4 Scout 17B 16E Instruct · 109B · 7.8%Gemma 3 4B IT · 4B · 7.5%Qwen2.5 32B Instruct · 33B · 7.4%Llama 3.3 70B Instruct · 71B · 5.1%Llama 3.1 Tulu 3 70B DPO · 71B · 4.4%Meta Llama 3 70B Instruct · 71B · 4.3%Llama 3.1 70B Instruct · 71B · 3.6%Granite 4.0 Micro · 3B · 2.8%Llama 3.2 90B Vision Instruct · 89B · 2.6%Hermes 2 Theta Llama 3 70B · 71B · 2.5%Qwen2.5 7B Instruct · 8B · 2.5%Meta Llama 3 8B Instruct · 8B · 1.9%Llama 3.1 8B Instruct · 8B · 1.7%Gemma 2 27B IT · 27B · 1.4%Deepseek Llm 67B Chat · 67B · 0.8%Gemma 2 9B IT · 9B · 0.6%Llama 3.2 1B Instruct · 1B · 0.6%Mistral 7B Instruct v0.3 · 7B · 0.3%Llama 2 70B Chat HF · 69B · 0.0%Gemma 3 1B IT · 1000M · 1.1%Gemma 3 1B ITDeepSeek R1 Distill Qwen 1.5B · 2B · 21.4%DeepSeek R1 Distill Q…Qwen3 4B Instruct 2507 · 4B · 52.2%Qwen3.5 4B · 5B · 55.8%Qwen3.5 4BQwen3 8B · 8B · 56.1%Qwen3 8BQwen3.5 9B · 10B · 61.7%Qwen3.5 9BQwen3 14B · 15B · 66.4%Qwen3 14BGemma 4 26B A4B IT · 26B · 82.2%Gemma 4 26B A4B ITQwen3.6 27B · 28B · 91.1%Qwen3.6 27BDeepSeek V4 Flash 0731 · 304B · 94.4%DeepSeek V4 Flash 0731Kimi K2.6 · 1T · 96.1%DeepSeek V4 Pro · 1.6T · 96.7%DeepSeek V4 ProDeepSeek V4 Pro 0813 · 1.7T · 98.6%DeepSeek V4 Pro 0813
Each dot is a model. Up = higher score, left = smaller (easier to run locally). The dashed line marks the efficiency frontier — the best score you can get at each size or smaller.
  • Gemma 3 1B IT, 1000M, score 1.1% — on the efficiency frontier (best score at its size or smaller).
  • DeepSeek R1 Distill Qwen 1.5B, 2B, score 21.4% — on the efficiency frontier (best score at its size or smaller).
  • Qwen3 4B Instruct 2507, 4B, score 52.2% — on the efficiency frontier (best score at its size or smaller).
  • Qwen3.5 4B, 5B, score 55.8% — on the efficiency frontier (best score at its size or smaller).
  • Qwen3 8B, 8B, score 56.1% — on the efficiency frontier (best score at its size or smaller).
  • Qwen3.5 9B, 10B, score 61.7% — on the efficiency frontier (best score at its size or smaller).
  • Qwen3 14B, 15B, score 66.4% — on the efficiency frontier (best score at its size or smaller).
  • Gemma 4 26B A4B IT, 26B, score 82.2% — on the efficiency frontier (best score at its size or smaller).
  • Qwen3.6 27B, 28B, score 91.1% — on the efficiency frontier (best score at its size or smaller).
  • DeepSeek V4 Flash 0731, 304B, score 94.4% — on the efficiency frontier (best score at its size or smaller).
  • Kimi K2.6, 1T, score 96.1% — on the efficiency frontier (best score at its size or smaller).
  • DeepSeek V4 Pro, 1.6T, score 96.7% — on the efficiency frontier (best score at its size or smaller).
  • DeepSeek V4 Pro 0813, 1.7T, score 98.6% — on the efficiency frontier (best score at its size or smaller).

AIME 2024/2025: frequently asked questions

What is the best open LLM on AIME 2024/2025?
DeepSeek V4 Pro 0813 is the top open model on AIME 2024/2025, scoring 98.6%. Among all models tested — including proprietary ones — it ranks #12. The top model overall is GPT 6 Astra Max (OpenAI) at 100.0%.
What's the best AIME 2024/2025 model you can run on a 24 GB GPU?
Qwen3.6 27B is the highest-scoring open model that fits in 24 GB at 4-bit quantization (about 15 GB), scoring 91.1% on AIME 2024/2025.
What's the best AIME 2024/2025 model you can run on a 12 GB GPU?
Qwen3 14B is the highest-scoring open model that fits in 12 GB at 4-bit quantization (about 8 GB), scoring 66.4% on AIME 2024/2025.
Can open models match proprietary models on AIME 2024/2025?
Not quite on AIME 2024/2025: the strongest proprietary model (GPT 6 Astra Max) scores 100.0%, ahead of the best open model (DeepSeek V4 Pro 0813) at 98.6% — but you can run the open one yourself.

Scores aggregated from epoch. llmrun does not run this benchmark — see the source for methodology, or the about benchmarks for what it measures.

AIME 2024/2025 Leaderboard — LLM Scores | llmrun