Best open LLMs for reasoning

Open models ranked by reasoning, from hard multi-step reasoning boards such as ARC-AGI and HLE. The score shown is the overall llmrun Score; the highlighted column is the reasoning skill score used for the order.

40 benchmarks4 sourcesUpdated 3 Oct 2026Reference scale 2026-Q4How the score is built

Models ranked by reasoning llmrun Score
#Modelllmrun ScoreCodingAgentsMathScienceReasoningBenchmarksVRAM
1DeepSeek V4 Pro 0813991 GB · 21 benchmarks63–71466521991 GB
2Kimi K31674 GB · 26 benchmarks7361815263261674 GB
3DeepSeek V4 Flash 0731183 GB · 21 benchmarks59–71456221183 GB
4GLM 5.3456 GB · 21 benchmarks676872495821456 GB
5DeepSeek V4.1 Flash458 GB · 10 benchmarks54–76455710458 GB
6Inkling Small160 GB · 12 benchmarks––57415412160 GB
7Inkling572 GB · 21 benchmarks43–52405321572 GB
8GLM 5.2456 GB · 24 benchmarks586068475224456 GB
9Kimi K2.6620 GB · 18 benchmarks57–63435118620 GB
10Qwen3.8 27B17.4 GB · 9 benchmarks51–573750917.4 GB
11Kimi K2.7 Code620 GB · 19 benchmarks534661405019620 GB
12DeepSeek V4 Pro960 GB · 22 benchmarks49–57444922960 GB
13NVIDIA Nemotron 3 Ultra 550B A55B BF16370 GB · 14 benchmarks43–52374614370 GB
14GLM 5.1457 GB · 13 benchmarks535457404613457 GB
15Qwen3.5 397B A17B266 GB · 9 benchmarks––52–469266 GB
16Kimi K2.5620 GB · 20 benchmarks4937–404520620 GB
17MiniMax M2.5138 GB · 11 benchmarks4737––4511138 GB
18DeepSeek V3.2 Exp415 GB · 10 benchmarks46––304410415 GB
19Qwen3.5 122B A10B82.6 GB · 6 benchmarks–––2643682.6 GB
20DeepSeek V3.2415 GB · 13 benchmarks3536––4313415 GB
21DeepSeek V4 Flash175 GB · 10 benchmarks46––364310175 GB
22MiMo V2.5 Pro615 GB · 6 benchmarks54––38426615 GB
23Qwen3.5 27B17.4 GB · 6 benchmarks38–––41617.4 GB
24Gemma 4 31B IT20.4 GB · 11 benchmarks47––32401120.4 GB
25GLM 5457 GB · 15 benchmarks4945––4015457 GB
26DeepSeek V3.1 Terminus415 GB · 5 benchmarks47––31395415 GB
27MiniMax M3257 GB · 19 benchmarks524046413819257 GB
28Qwen3.6 35B A3B22.0 GB · 12 benchmarks44–4935381222.0 GB
29Qwen3.6 27B17.4 GB · 12 benchmarks42–5336371217.4 GB
30GLM 5.3 Flash195 GB · 17 benchmarks526362443717195 GB
31Qwen3 235B A22B Thinking 2507142 GB · 13 benchmarks42––343713142 GB
32Qwen3.5 35B A3B22.0 GB · 8 benchmarks–––3437822.0 GB
33DeepSeek V3.1415 GB · 7 benchmarks––––367415 GB
34GLM 4.7216 GB · 13 benchmarks413251363513216 GB
35Gemma 4 26B A4B IT16.1 GB · 9 benchmarks43––3033916.1 GB
36DeepSeek R1 0528415 GB · 11 benchmarks46–44–3211415 GB
37Qwen3 235B A22B142 GB · 10 benchmarks41–––3110142 GB
38Qwen3 235B A22B Instruct 2507142 GB · 8 benchmarks41–––318142 GB
39Qwen3.5 9B6.4 GB · 9 benchmarks–––323096.4 GB
40GPT OSS 120B70.5 GB · 20 benchmarks3714–31302070.5 GB
41DeepSeek R1415 GB · 14 benchmarks39–40282914415 GB
42Mistral Small 4 119B 260373.2 GB · 3 benchmarks—––––29373.2 GB
43Qwen3 30B A3B Thinking 250718.7 GB · 8 benchmarks–––2827818.7 GB
44Qwen3 30B A3B Instruct 250718.7 GB · 5 benchmarks––––26518.7 GB
45Qwen3 32B20.3 GB · 9 benchmarks33––2626920.3 GB
46GPT OSS 20B12.9 GB · 11 benchmarks40––24261112.9 GB
47Kimi K2 Instruct620 GB · 10 benchmarks4128––2510620 GB
48Mistral Large 3 675B Instruct 2512446 GB · 5 benchmarks33––26255446 GB
49Magistral Small 250614.9 GB · 4 benchmarks––––24414.9 GB
50DeepSeek v3 0324415 GB · 13 benchmarks39–33272413415 GB
51Qwen3 14B9.5 GB · 8 benchmarks–––252489.5 GB
52Qwen2.5 72B Instruct44.6 GB · 9 benchmarks––25–23944.6 GB
53Llama 4 Maverick 17B 128E Instruct265 GB · 17 benchmarks25–29262317265 GB
54Llama 3.1 405B Instruct268 GB · 8 benchmarks––22–228268 GB
55Magistral Small 250915.1 GB · 6 benchmarks–––1822615.1 GB
56Mistral Large Instruct 241174.6 GB · 6 benchmarks––22–22674.6 GB
57C4ai Command A 03 202573.3 GB · 3 benchmarks—––––22373.3 GB
58Qwen3 30B A3B18.7 GB · 7 benchmarks––––22718.7 GB
59Llama 3.1 70B Instruct46.6 GB · 8 benchmarks––18–22846.6 GB
60Mistral Large Instruct 240780.9 GB · 7 benchmarks––20–22780.9 GB
61Mistral Small 3.2 24B Instruct 250615.1 GB · 6 benchmarks–––1822615.1 GB
62Llama 3.3 70B Instruct46.6 GB · 13 benchmarks22–1917211346.6 GB
63Mistral Small 3.1 24B Instruct 250315.1 GB · 7 benchmarks––211721715.1 GB
64DeepSeek v3415 GB · 10 benchmarks37–26222110415 GB
65Qwen3 8B5.5 GB · 8 benchmarks–––212185.5 GB
66Llama 4 Scout 17B 16E Instruct71.7 GB · 12 benchmarks––2418201271.7 GB
67Mixtral 8x22B Instruct v0.185.2 GB · 6 benchmarks––––19685.2 GB
68C4ai Command R Plus 08 202468.5 GB · 3 benchmarks—––––18368.5 GB
69Gemma 3 27B IT18.1 GB · 11 benchmarks18–2916171118.1 GB
70Llama 3.1 8B Instruct5.3 GB · 11 benchmarks14–14716115.3 GB
71Mixtral 8x7B Instruct v0.128.6 GB · 5 benchmarks––––16528.6 GB
72Gemma 3 4B IT2.8 GB · 6 benchmarks––––1662.8 GB
73Gemma 3 12B IT8.0 GB · 8 benchmarks–––121588.0 GB
74Qwen2.5 7B Instruct5.0 GB · 6 benchmarks––––1565.0 GB
75C4ai Command R 08 202421.3 GB · 2 benchmarks—––––14221.3 GB
76Meta Llama 3 8B Instruct5.3 GB · 7 benchmarks––10–1475.3 GB
77Llama 2 13B Chat HF8.6 GB · 2 benchmarks—––––1228.6 GB
78Llama 2 70B Chat HF45.5 GB · 5 benchmarks––9–10545.5 GB

The range after each score is a 90% interval: given the boards a model has results on, its true score is very likely inside it. Short ranges mean many agreeing results; long ranges mean few or conflicting ones.

Why some models are missing. A model gets a score only with results on at least 4 benchmarks across at least 2 categories, one of them reasoning or coding. Skill scores need at least 2 benchmarks in that skill; otherwise they show as –.

VRAM is llmrun's estimate at Q4_K_M (or the smallest quantization we track) for the weights plus a working context. Proprietary models can't run locally, so the VRAM filter hides them.

How the llmrun Score is built · llmrun does not run these benchmarks; scores are aggregated from public sources.