Math
AIME 2024/2025 Leaderboard
AIME (American Invitational Mathematics Examination) is a prestigious high-school competition of hard, integer-answer problems. It's a widely-cited yardstick for multi-step mathematical reasoning, here on the 2024–2025 papers.
Source: epoch74 open models ranked+196 proprietaryData through Sep 2026
Open models ranked on AIME 2024/2025
# shows rank among open models / rank overall (including proprietary).
| # | Model | Score |
|---|---|---|
| 1 / 12 | DeepSeek V4 Pro 0813 · 1650.5B | 98.6% |
| 2 / 23 | Kimi K3 · 2779.9B | 97.2% |
| 3 / 24 | DeepSeek V4 Pro · 1598.8B | 96.7% |
| 4 / 26 | Kimi K2.6 · 1026.9B | 96.1% |
| 5 / 34 | Kimi K2.7 Code · 1026.9B | 95.6% |
| 6 / 39 | DeepSeek V4 Flash 0731 · 304.2B | 94.4% |
| 7 / 41 | GLM 5.3 Flash · 321.3B | 93.9% |
| 8 / 44 | GLM 5.1 · 753.9B | 93.3% |
| 9 / 51 | Kimi K2.5 · 1026.9B | 92.2% |
| 10 / 55 | GLM 5.3 · 753.3B | 91.1% |
| 11 / 56 | Qwen3.6 27B · 27.8B | 91.1% |
| 12 / 58 | Inkling Small · 266.0B | 90.0% |
| 13 / 63 | GPT OSS 120B · 116.8B | 88.9% |
| 14 / 64 | Inkling · 952.4B | 88.9% |
| 15 / 65 | Qwen3.5 397B A17B · 403.4B | 88.9% |
| 16 / 74 | Qwen3 235B A22B Thinking 2507 · 235.1B | 86.7% |
| 17 / 76 | Qwen3.6 35B A3B · 36.0B | 86.7% |
| 18 / 78 | GLM 5.2 · 753.3B | 86.4% |
| 19 / 91 | GLM 4.7 · 358.3B | 83.3% |
| 20 / 92 | Kimi K2 Thinking · 1026.4B | 83.1% |
| 21 / 95 | Gemma 4 26B A4B IT · 25.8B | 82.2% |
| 22 / 103 | GLM 5 · 753.9B | 80.0% |
| 23 / 114 | Gemma 4 31B IT · 31.3B | 73.3% |
| 24 / 124 | MiniMax M3 · 427.0B | 71.1% |
| 25 / 126 | Qwen3 30B A3B Thinking 2507 · 30.5B | 70.3% |
| 26 / 133 | Seed OSS 36B Instruct · 36.2B | 67.5% |
| 27 / 134 | Qwen3 32B · 32.8B | 66.9% |
| 28 / 137 | DeepSeek R1 0528 · 684.5B | 66.4% |
| 29 / 138 | Qwen3 14B · 14.8B | 66.4% |
| 30 / 139 | GPT OSS 20B · 20.9B | 65.3% |
| 31 / 144 | Qwen3 30B A3B · 30.5B | 62.8% |
| 32 / 147 | Qwen3 30B A3B Instruct 2507 · 30.5B | 62.2% |
| 33 / 148 | Qwen3.5 9B · 9.7B | 61.7% |
| 34 / 152 | QwQ 32B · 32.8B | 59.2% |
| 35 / 153 | GLM 4.7 Flash · 31.2B | 58.3% |
| 36 / 159 | Qwen3 8B · 8.2B | 56.1% |
| 37 / 160 | Qwen3.5 4B · 4.7B | 55.8% |
| 38 / 161 | DeepSeek R1 Distill Qwen 32B · 32.8B | 55.6% |
| 39 / 167 | DeepSeek R1 · 684.5B | 53.3% |
| 40 / 170 | Qwen3 4B Instruct 2507 · 4.0B | 52.2% |
| 41 / 171 | DeepSeek R1 Distill Llama 70B · 70.6B | 51.4% |
| 42 / 173 | DeepSeek R1 Distill Qwen 14B · 14.8B | 50.6% |
| 43 / 185 | DeepSeek R1 0528 Qwen3 8B · 8.2B | 43.9% |
| 44 / 190 | DeepSeek v3 0324 · 684.5B | 37.8% |
| 45 / 201 | Magistral Small 2506 · 23.6B | 30.0% |
| 46 / 207 | Gemma 3 27B IT · 27.4B | 22.5% |
| 47 / 209 | DeepSeek R1 Distill Qwen 1.5B · 1.8B | 21.4% |
| 48 / 210 | Llama 4 Maverick 17B 128E Instruct · 401.6B | 20.6% |
| 49 / 212 | Gemma 3 12B IT · 12.2B | 16.7% |
| 50 / 215 | DeepSeek v3 · 684.5B | 15.8% |
| 51 / 216 | Phi 4 · 14.7B | 13.8% |
| 52 / 218 | Llama 3.1 405B Instruct · 405.9B | 9.7% |
| 53 / 221 | Qwen2.5 72B Instruct · 72.7B | 8.1% |
| 54 / 222 | Qwen3 1.7B · 2.0B | 8.1% |
| 55 / 223 | Llama 4 Scout 17B 16E Instruct · 108.6B | 7.8% |
| 56 / 225 | Gemma 3 4B IT · 4.3B | 7.5% |
| 57 / 226 | Qwen2.5 32B Instruct · 32.8B | 7.4% |
| 58 / 238 | Llama 3.3 70B Instruct · 70.6B | 5.1% |
| 59 / 241 | Llama 3.1 Tulu 3 70B DPO · 70.6B | 4.4% |
| 60 / 243 | Meta Llama 3 70B Instruct · 70.6B | 4.3% |
| 61 / 246 | Llama 3.1 70B Instruct · 70.6B | 3.6% |
| 62 / 247 | Granite 4.0 Micro · 3.4B | 2.8% |
| 63 / 248 | Llama 3.2 90B Vision Instruct · 88.6B | 2.6% |
| 64 / 251 | Hermes 2 Theta Llama 3 70B · 70.6B | 2.5% |
| 65 / 252 | Qwen2.5 7B Instruct · 7.6B | 2.5% |
| 66 / 255 | Meta Llama 3 8B Instruct · 8.0B | 1.9% |
| 67 / 258 | Llama 3.1 8B Instruct · 8.0B | 1.7% |
| 68 / 259 | Gemma 2 27B IT · 27.2B | 1.4% |
| 69 / 261 | Gemma 3 1B IT · 1000M | 1.1% |
| 70 / 263 | Deepseek Llm 67B Chat · 67B | 0.8% |
| 71 / 265 | Gemma 2 9B IT · 9.2B | 0.6% |
| 72 / 267 | Llama 3.2 1B Instruct · 1.2B | 0.6% |
| 73 / 269 | Mistral 7B Instruct v0.3 · 7.2B | 0.3% |
| 74 / 270 | Llama 2 70B Chat HF · 69.0B | 0.0% |
Score vs model size
Which models give the most quality for their size — the ones worth running locally.
- Gemma 3 1B IT, 1000M, score 1.1% — on the efficiency frontier (best score at its size or smaller).
- DeepSeek R1 Distill Qwen 1.5B, 2B, score 21.4% — on the efficiency frontier (best score at its size or smaller).
- Qwen3 4B Instruct 2507, 4B, score 52.2% — on the efficiency frontier (best score at its size or smaller).
- Qwen3.5 4B, 5B, score 55.8% — on the efficiency frontier (best score at its size or smaller).
- Qwen3 8B, 8B, score 56.1% — on the efficiency frontier (best score at its size or smaller).
- Qwen3.5 9B, 10B, score 61.7% — on the efficiency frontier (best score at its size or smaller).
- Qwen3 14B, 15B, score 66.4% — on the efficiency frontier (best score at its size or smaller).
- Gemma 4 26B A4B IT, 26B, score 82.2% — on the efficiency frontier (best score at its size or smaller).
- Qwen3.6 27B, 28B, score 91.1% — on the efficiency frontier (best score at its size or smaller).
- DeepSeek V4 Flash 0731, 304B, score 94.4% — on the efficiency frontier (best score at its size or smaller).
- Kimi K2.6, 1T, score 96.1% — on the efficiency frontier (best score at its size or smaller).
- DeepSeek V4 Pro, 1.6T, score 96.7% — on the efficiency frontier (best score at its size or smaller).
- DeepSeek V4 Pro 0813, 1.7T, score 98.6% — on the efficiency frontier (best score at its size or smaller).
AIME 2024/2025: frequently asked questions
- What is the best open LLM on AIME 2024/2025?
- DeepSeek V4 Pro 0813 is the top open model on AIME 2024/2025, scoring 98.6%. Among all models tested — including proprietary ones — it ranks #12. The top model overall is GPT 6 Astra Max (OpenAI) at 100.0%.
- What's the best AIME 2024/2025 model you can run on a 24 GB GPU?
- Qwen3.6 27B is the highest-scoring open model that fits in 24 GB at 4-bit quantization (about 15 GB), scoring 91.1% on AIME 2024/2025.
- What's the best AIME 2024/2025 model you can run on a 12 GB GPU?
- Qwen3 14B is the highest-scoring open model that fits in 12 GB at 4-bit quantization (about 8 GB), scoring 66.4% on AIME 2024/2025.
- Can open models match proprietary models on AIME 2024/2025?
- Not quite on AIME 2024/2025: the strongest proprietary model (GPT 6 Astra Max) scores 100.0%, ahead of the best open model (DeepSeek V4 Pro 0813) at 98.6% — but you can run the open one yourself.
Scores aggregated from epoch. llmrun does not run this benchmark — see the source for methodology, or the about benchmarks for what it measures.