Best open LLMs for math

Open models ranked by mathematics, from competition and research math boards such as AIME and FrontierMath. The score shown is the overall llmrun Score; the highlighted column is the mathematics skill score used for the order.

40 benchmarks4 sourcesUpdated 3 Oct 2026Reference scale 2026-Q4How the score is built

Models ranked by math llmrun Score
#Modelllmrun ScoreCodingAgentsMathScienceReasoningBenchmarksVRAM
1Kimi K31674 GB · 26 benchmarks7361815263261674 GB
2DeepSeek V4.1 Flash458 GB · 10 benchmarks54–76455710458 GB
3GLM 5.3456 GB · 21 benchmarks676872495821456 GB
4DeepSeek V4 Pro 0813991 GB · 21 benchmarks63–71466521991 GB
5DeepSeek V4 Flash 0731183 GB · 21 benchmarks59–71456221183 GB
6GLM 5.2456 GB · 24 benchmarks586068475224456 GB
7Kimi K2.6620 GB · 18 benchmarks57–63435118620 GB
8GLM 5.3 Flash195 GB · 17 benchmarks526362443717195 GB
9Kimi K2.7 Code620 GB · 19 benchmarks534661405019620 GB
10DeepSeek V4 Pro960 GB · 22 benchmarks49–57444922960 GB
11Qwen3.8 27B17.4 GB · 9 benchmarks51–573750917.4 GB
12GLM 5.1457 GB · 13 benchmarks535457404613457 GB
13Inkling Small160 GB · 12 benchmarks––57415412160 GB
14Qwen3.6 27B17.4 GB · 12 benchmarks42–5336371217.4 GB
15Qwen3.5 397B A17B266 GB · 9 benchmarks––52–469266 GB
16Inkling572 GB · 21 benchmarks43–52405321572 GB
17NVIDIA Nemotron 3 Ultra 550B A55B BF16370 GB · 14 benchmarks43–52374614370 GB
18GLM 4.7216 GB · 13 benchmarks413251363513216 GB
19Qwen3.6 35B A3B22.0 GB · 12 benchmarks44–4935381222.0 GB
20MiniMax M3257 GB · 19 benchmarks524046413819257 GB
21DeepSeek R1 0528415 GB · 11 benchmarks46–44–3211415 GB
22DeepSeek R1415 GB · 14 benchmarks39–40282914415 GB
23DeepSeek R1 Distill Llama 70B43.3 GB · 3 benchmarks—––39––343.3 GB
24DeepSeek R1 Distill Qwen 14B9.6 GB · 4 benchmarks––38––49.6 GB
25DeepSeek v3 0324415 GB · 13 benchmarks39–33272413415 GB
26Gemma 3 27B IT18.1 GB · 11 benchmarks18–2916171118.1 GB
27Llama 4 Maverick 17B 128E Instruct265 GB · 17 benchmarks25–29262317265 GB
28DeepSeek v3415 GB · 10 benchmarks37–26222110415 GB
29Phi 49.5 GB · 6 benchmarks––26––69.5 GB
30Qwen2.5 72B Instruct44.6 GB · 9 benchmarks––25–23944.6 GB
31Llama 4 Scout 17B 16E Instruct71.7 GB · 12 benchmarks––2418201271.7 GB
32Qwen2.5 32B Instruct20.5 GB · 5 benchmarks––23––520.5 GB
33Llama 3.1 405B Instruct268 GB · 8 benchmarks––22–228268 GB
34Mistral Large Instruct 241174.6 GB · 6 benchmarks––22–22674.6 GB
35Mistral Small 3.1 24B Instruct 250315.1 GB · 7 benchmarks––211721715.1 GB
36Mistral Large Instruct 240780.9 GB · 7 benchmarks––20–22780.9 GB
37Mistral Small 24B Instruct 250114.9 GB · 4 benchmarks––20––414.9 GB
38Llama 3.1 Tulu 3 70B DPO43.3 GB · 3 benchmarks—––19––343.3 GB
39Llama 3.3 70B Instruct46.6 GB · 13 benchmarks22–1917211346.6 GB
40Llama 3.2 90B Vision Instruct58.5 GB · 4 benchmarks––18––458.5 GB
41Llama 3.1 70B Instruct46.6 GB · 8 benchmarks––18–22846.6 GB
42Gemma 2 27B IT18.0 GB · 4 benchmarks––16––418.0 GB
43Meta Llama 3 70B Instruct46.6 GB · 5 benchmarks––14––546.6 GB
44Hermes 2 Theta Llama 3 70B43.3 GB · 3 benchmarks—––14––343.3 GB
45Llama 3.1 8B Instruct5.3 GB · 11 benchmarks14–14716115.3 GB
46Gemma 2 9B IT6.1 GB · 4 benchmarks––14––46.1 GB
47Deepseek Llm 67B Chat41.3 GB · 4 benchmarks––10––441.3 GB
48Meta Llama 3 8B Instruct5.3 GB · 7 benchmarks––10–1475.3 GB
49Mistral 7B Instruct v0.34.9 GB · 4 benchmarks––10––44.9 GB
50Llama 2 70B Chat HF45.5 GB · 5 benchmarks––9–10545.5 GB

The range after each score is a 90% interval: given the boards a model has results on, its true score is very likely inside it. Short ranges mean many agreeing results; long ranges mean few or conflicting ones.

Why some models are missing. A model gets a score only with results on at least 4 benchmarks across at least 2 categories, one of them reasoning or coding. Skill scores need at least 2 benchmarks in that skill; otherwise they show as –.

VRAM is llmrun's estimate at Q4_K_M (or the smallest quantization we track) for the weights plus a working context. Proprietary models can't run locally, so the VRAM filter hides them.

How the llmrun Score is built · llmrun does not run these benchmarks; scores are aggregated from public sources.