Best open LLMs for math
Open models ranked by mathematics, from competition and research math boards such as AIME and FrontierMath. The score shown is the overall llmrun Score; the highlighted column is the mathematics skill score used for the order.
40 benchmarks4 sourcesUpdated 3 Oct 2026Reference scale 2026-Q4How the score is built
| # | Model | llmrun Score | Coding | Agents | Math | Science | Reasoning | Benchmarks | VRAM |
|---|---|---|---|---|---|---|---|---|---|
| 1 | Kimi K31674 GB · 26 benchmarks | 73 | 61 | 81 | 52 | 63 | 26 | 1674 GB | |
| 2 | DeepSeek V4.1 Flash458 GB · 10 benchmarks | 54 | – | 76 | 45 | 57 | 10 | 458 GB | |
| 3 | GLM 5.3456 GB · 21 benchmarks | 67 | 68 | 72 | 49 | 58 | 21 | 456 GB | |
| 4 | DeepSeek V4 Pro 0813991 GB · 21 benchmarks | 63 | – | 71 | 46 | 65 | 21 | 991 GB | |
| 5 | DeepSeek V4 Flash 0731183 GB · 21 benchmarks | 59 | – | 71 | 45 | 62 | 21 | 183 GB | |
| 6 | GLM 5.2456 GB · 24 benchmarks | 58 | 60 | 68 | 47 | 52 | 24 | 456 GB | |
| 7 | Kimi K2.6620 GB · 18 benchmarks | 57 | – | 63 | 43 | 51 | 18 | 620 GB | |
| 8 | GLM 5.3 Flash195 GB · 17 benchmarks | 52 | 63 | 62 | 44 | 37 | 17 | 195 GB | |
| 9 | Kimi K2.7 Code620 GB · 19 benchmarks | 53 | 46 | 61 | 40 | 50 | 19 | 620 GB | |
| 10 | DeepSeek V4 Pro960 GB · 22 benchmarks | 49 | – | 57 | 44 | 49 | 22 | 960 GB | |
| 11 | Qwen3.8 27B17.4 GB · 9 benchmarks | 51 | – | 57 | 37 | 50 | 9 | 17.4 GB | |
| 12 | GLM 5.1457 GB · 13 benchmarks | 53 | 54 | 57 | 40 | 46 | 13 | 457 GB | |
| 13 | Inkling Small160 GB · 12 benchmarks | – | – | 57 | 41 | 54 | 12 | 160 GB | |
| 14 | Qwen3.6 27B17.4 GB · 12 benchmarks | 42 | – | 53 | 36 | 37 | 12 | 17.4 GB | |
| 15 | Qwen3.5 397B A17B266 GB · 9 benchmarks | – | – | 52 | – | 46 | 9 | 266 GB | |
| 16 | Inkling572 GB · 21 benchmarks | 43 | – | 52 | 40 | 53 | 21 | 572 GB | |
| 17 | NVIDIA Nemotron 3 Ultra 550B A55B BF16370 GB · 14 benchmarks | 43 | – | 52 | 37 | 46 | 14 | 370 GB | |
| 18 | GLM 4.7216 GB · 13 benchmarks | 41 | 32 | 51 | 36 | 35 | 13 | 216 GB | |
| 19 | Qwen3.6 35B A3B22.0 GB · 12 benchmarks | 44 | – | 49 | 35 | 38 | 12 | 22.0 GB | |
| 20 | MiniMax M3257 GB · 19 benchmarks | 52 | 40 | 46 | 41 | 38 | 19 | 257 GB | |
| 21 | DeepSeek R1 0528415 GB · 11 benchmarks | 46 | – | 44 | – | 32 | 11 | 415 GB | |
| 22 | DeepSeek R1415 GB · 14 benchmarks | 39 | – | 40 | 28 | 29 | 14 | 415 GB | |
| 23 | DeepSeek R1 Distill Llama 70B43.3 GB · 3 benchmarks | — | – | – | 39 | – | – | 3 | 43.3 GB |
| 24 | DeepSeek R1 Distill Qwen 14B9.6 GB · 4 benchmarks | – | – | 38 | – | – | 4 | 9.6 GB | |
| 25 | DeepSeek v3 0324415 GB · 13 benchmarks | 39 | – | 33 | 27 | 24 | 13 | 415 GB | |
| 26 | Gemma 3 27B IT18.1 GB · 11 benchmarks | 18 | – | 29 | 16 | 17 | 11 | 18.1 GB | |
| 27 | Llama 4 Maverick 17B 128E Instruct265 GB · 17 benchmarks | 25 | – | 29 | 26 | 23 | 17 | 265 GB | |
| 28 | DeepSeek v3415 GB · 10 benchmarks | 37 | – | 26 | 22 | 21 | 10 | 415 GB | |
| 29 | Phi 49.5 GB · 6 benchmarks | – | – | 26 | – | – | 6 | 9.5 GB | |
| 30 | Qwen2.5 72B Instruct44.6 GB · 9 benchmarks | – | – | 25 | – | 23 | 9 | 44.6 GB | |
| 31 | Llama 4 Scout 17B 16E Instruct71.7 GB · 12 benchmarks | – | – | 24 | 18 | 20 | 12 | 71.7 GB | |
| 32 | Qwen2.5 32B Instruct20.5 GB · 5 benchmarks | – | – | 23 | – | – | 5 | 20.5 GB | |
| 33 | Llama 3.1 405B Instruct268 GB · 8 benchmarks | – | – | 22 | – | 22 | 8 | 268 GB | |
| 34 | Mistral Large Instruct 241174.6 GB · 6 benchmarks | – | – | 22 | – | 22 | 6 | 74.6 GB | |
| 35 | Mistral Small 3.1 24B Instruct 250315.1 GB · 7 benchmarks | – | – | 21 | 17 | 21 | 7 | 15.1 GB | |
| 36 | Mistral Large Instruct 240780.9 GB · 7 benchmarks | – | – | 20 | – | 22 | 7 | 80.9 GB | |
| 37 | Mistral Small 24B Instruct 250114.9 GB · 4 benchmarks | – | – | 20 | – | – | 4 | 14.9 GB | |
| 38 | Llama 3.1 Tulu 3 70B DPO43.3 GB · 3 benchmarks | — | – | – | 19 | – | – | 3 | 43.3 GB |
| 39 | Llama 3.3 70B Instruct46.6 GB · 13 benchmarks | 22 | – | 19 | 17 | 21 | 13 | 46.6 GB | |
| 40 | Llama 3.2 90B Vision Instruct58.5 GB · 4 benchmarks | – | – | 18 | – | – | 4 | 58.5 GB | |
| 41 | Llama 3.1 70B Instruct46.6 GB · 8 benchmarks | – | – | 18 | – | 22 | 8 | 46.6 GB | |
| 42 | Gemma 2 27B IT18.0 GB · 4 benchmarks | – | – | 16 | – | – | 4 | 18.0 GB | |
| 43 | Meta Llama 3 70B Instruct46.6 GB · 5 benchmarks | – | – | 14 | – | – | 5 | 46.6 GB | |
| 44 | Hermes 2 Theta Llama 3 70B43.3 GB · 3 benchmarks | — | – | – | 14 | – | – | 3 | 43.3 GB |
| 45 | Llama 3.1 8B Instruct5.3 GB · 11 benchmarks | 14 | – | 14 | 7 | 16 | 11 | 5.3 GB | |
| 46 | Gemma 2 9B IT6.1 GB · 4 benchmarks | – | – | 14 | – | – | 4 | 6.1 GB | |
| 47 | Deepseek Llm 67B Chat41.3 GB · 4 benchmarks | – | – | 10 | – | – | 4 | 41.3 GB | |
| 48 | Meta Llama 3 8B Instruct5.3 GB · 7 benchmarks | – | – | 10 | – | 14 | 7 | 5.3 GB | |
| 49 | Mistral 7B Instruct v0.34.9 GB · 4 benchmarks | – | – | 10 | – | – | 4 | 4.9 GB | |
| 50 | Llama 2 70B Chat HF45.5 GB · 5 benchmarks | – | – | 9 | – | 10 | 5 | 45.5 GB |
The range after each score is a 90% interval: given the boards a model has results on, its true score is very likely inside it. Short ranges mean many agreeing results; long ranges mean few or conflicting ones.
Why some models are missing. A model gets a score only with results on at least 4 benchmarks across at least 2 categories, one of them reasoning or coding. Skill scores need at least 2 benchmarks in that skill; otherwise they show as –.
VRAM is llmrun's estimate at Q4_K_M (or the smallest quantization we track) for the weights plus a working context. Proprietary models can't run locally, so the VRAM filter hides them.
How the llmrun Score is built · llmrun does not run these benchmarks; scores are aggregated from public sources.