Best open LLMs for science

Open models ranked by scientific knowledge, from graduate-level science and knowledge boards such as GPQA. The score shown is the overall llmrun Score; the highlighted column is the scientific knowledge skill score used for the order.

40 benchmarks4 sourcesUpdated 3 Oct 2026Reference scale 2026-Q4How the score is built

Models ranked by science llmrun Score
#Modelllmrun ScoreCodingAgentsMathScienceReasoningBenchmarksVRAM
1Kimi K31674 GB · 26 benchmarks7361815263261674 GB
2GLM 5.3456 GB · 21 benchmarks676872495821456 GB
3GLM 5.2456 GB · 24 benchmarks586068475224456 GB
4DeepSeek V4 Pro 0813991 GB · 21 benchmarks63–71466521991 GB
5DeepSeek V4.1 Flash458 GB · 10 benchmarks54–76455710458 GB
6DeepSeek V4 Flash 0731183 GB · 21 benchmarks59–71456221183 GB
7GLM 5.3 Flash195 GB · 17 benchmarks526362443717195 GB
8DeepSeek V4 Pro960 GB · 22 benchmarks49–57444922960 GB
9Kimi K2.6620 GB · 18 benchmarks57–63435118620 GB
10Inkling Small160 GB · 12 benchmarks––57415412160 GB
11MiniMax M3257 GB · 19 benchmarks524046413819257 GB
12Kimi K2.7 Code620 GB · 19 benchmarks534661405019620 GB
13Kimi K2.5620 GB · 20 benchmarks4937–404520620 GB
14GLM 5.1457 GB · 13 benchmarks535457404613457 GB
15Inkling572 GB · 21 benchmarks43–52405321572 GB
16MiMo V2.5 Pro615 GB · 6 benchmarks54––38426615 GB
17Qwen3.8 27B17.4 GB · 9 benchmarks51–573750917.4 GB
18NVIDIA Nemotron 3 Ultra 550B A55B BF16370 GB · 14 benchmarks43–52374614370 GB
19DeepSeek V4 Flash175 GB · 10 benchmarks46––364310175 GB
20GLM 4.7216 GB · 13 benchmarks413251363513216 GB
21Qwen3.6 27B17.4 GB · 12 benchmarks42–5336371217.4 GB
22Qwen3.6 35B A3B22.0 GB · 12 benchmarks44–4935381222.0 GB
23MiniMax M2.7138 GB · 6 benchmarks41––35–6138 GB
24Qwen3 235B A22B Thinking 2507142 GB · 13 benchmarks42––343713142 GB
25Qwen3.5 35B A3B22.0 GB · 8 benchmarks–––3437822.0 GB
26MiMo V2.5187 GB · 4 benchmarks43––33–4187 GB
27Ring 2.6 1T677 GB · 3 benchmarks—40––33–3677 GB
28Gemma 4 31B IT20.4 GB · 11 benchmarks47––32401120.4 GB
29Qwen3.5 9B6.4 GB · 9 benchmarks–––323096.4 GB
30GPT OSS 120B70.5 GB · 20 benchmarks3714–31302070.5 GB
31DeepSeek V3.1 Terminus415 GB · 5 benchmarks47––31395415 GB
32Step 3.7 Flash133 GB · 3 benchmarks—46––30–3133 GB
33Cogito 671B V2.1407 GB · 2 benchmarks—–––30–2407 GB
34Gemma 4 26B A4B IT16.1 GB · 9 benchmarks43––3033916.1 GB
35DeepSeek V3.2 Exp415 GB · 10 benchmarks46––304410415 GB
36GLM 4.6215 GB · 6 benchmarks3631–29–6215 GB
37Qwen3 30B A3B Thinking 250718.7 GB · 8 benchmarks–––2827818.7 GB
38Command A Plus 05 2026 BF16132 GB · 2 benchmarks—–––28–2132 GB
39DeepSeek R1415 GB · 14 benchmarks39–40282914415 GB
40NVIDIA Nemotron 3 Super 120B A12B BF1674.7 GB · 4 benchmarks36––27–474.7 GB
41DeepSeek v3 0324415 GB · 13 benchmarks39–33272413415 GB
42Trinity Large Thinking263 GB · 3 benchmarks—32––26–3263 GB
43Mistral Large 3 675B Instruct 2512446 GB · 5 benchmarks33––26255446 GB
44Qwen3 32B20.3 GB · 9 benchmarks33––2626920.3 GB
45Llama 4 Maverick 17B 128E Instruct265 GB · 17 benchmarks25–29262317265 GB
46Qwen3.5 122B A10B82.6 GB · 6 benchmarks–––2643682.6 GB
47Qwen3 14B9.5 GB · 8 benchmarks–––252489.5 GB
48GPT OSS 20B12.9 GB · 11 benchmarks40––24261112.9 GB
49Qwen3 Coder Next48.2 GB · 3 benchmarks—35––23–348.2 GB
50DeepSeek v3415 GB · 10 benchmarks37–26222110415 GB
51Qwen3 8B5.5 GB · 8 benchmarks–––212185.5 GB
52Devstral Small 2 24B Instruct 251215.1 GB · 3 benchmarks—–––20–315.1 GB
53Llama 4 Scout 17B 16E Instruct71.7 GB · 12 benchmarks––2418201271.7 GB
54Magistral Small 250915.1 GB · 6 benchmarks–––1822615.1 GB
55Mistral Small 3.2 24B Instruct 250615.1 GB · 6 benchmarks–––1822615.1 GB
56MiMo v2 Flash186 GB · 3 benchmarks—40––17–3186 GB
57Granite 4.1 30B18.2 GB · 2 benchmarks—–––17–218.2 GB
58Mistral Small 3.1 24B Instruct 250315.1 GB · 7 benchmarks––211721715.1 GB
59Llama 3.3 70B Instruct46.6 GB · 13 benchmarks22–1917211346.6 GB
60Gemma 3 27B IT18.1 GB · 11 benchmarks18–2916171118.1 GB
61Gemma 3 12B IT8.0 GB · 8 benchmarks–––121588.0 GB
62Phi 4 Mini Instruct2.9 GB · 3 benchmarks—–––9–32.9 GB
63Llama 3.1 8B Instruct5.3 GB · 11 benchmarks14–14716115.3 GB

The range after each score is a 90% interval: given the boards a model has results on, its true score is very likely inside it. Short ranges mean many agreeing results; long ranges mean few or conflicting ones.

Why some models are missing. A model gets a score only with results on at least 4 benchmarks across at least 2 categories, one of them reasoning or coding. Skill scores need at least 2 benchmarks in that skill; otherwise they show as –.

VRAM is llmrun's estimate at Q4_K_M (or the smallest quantization we track) for the weights plus a working context. Proprietary models can't run locally, so the VRAM filter hides them.

How the llmrun Score is built · llmrun does not run these benchmarks; scores are aggregated from public sources.