Best LLMs for science: open and proprietary

Open and proprietary models together on one 0–100 scale (scientific knowledge). Switch to local models to see only what you can download and run.

40 benchmarks4 sourcesUpdated 3 Oct 2026Reference scale 2026-Q4How the score is built

Models ranked by science llmrun Score
#Modelllmrun ScoreCodingAgentsMathScienceReasoningBenchmarksVRAM
121Gemini 1.5 Pro 0029 benchmarks––2822219–
122Claude 3.5 Sonnet (Oct 22, 2024)11 benchmarks382623212911–
123Qwen3 8Bopen5.5 GB · 8 benchmarks–––212185.5 GB
124Devstral Small 2 24B Instruct 2512open15.1 GB · 3 benchmarks—–––20–315.1 GB
125Llama 4 Scout 17B 16E Instructopen71.7 GB · 12 benchmarks––2418201271.7 GB
126Magistral Small 2509open15.1 GB · 6 benchmarks–––1822615.1 GB
127Mistral Small 3.2 24B Instruct 2506open15.1 GB · 6 benchmarks–––1822615.1 GB
128GPT 4.1 Nano (Apr 14, 2025)13 benchmarks21–30171713–
129MiMo v2 Flashopen186 GB · 3 benchmarks—40––17–3186 GB
130Granite 4.1 30Bopen18.2 GB · 2 benchmarks—–––17–218.2 GB
131GPT 4o (Nov 20, 2024)10 benchmarks25–21172410–
132Mistral Small 3.1 24B Instruct 2503open15.1 GB · 7 benchmarks––211721715.1 GB
133Llama 3.3 70B Instructopen46.6 GB · 13 benchmarks22–1917211346.6 GB
134Gemma 3 27B ITopen18.1 GB · 11 benchmarks18–2916171118.1 GB
135Solar Pro 32 benchmarks—–––16–2–
136Claude 3.5 Haiku (Oct 22, 2024)9 benchmarks29–2013–9–
137Gemma 3 12B ITopen8.0 GB · 8 benchmarks–––121588.0 GB
138Phi 4 Mini Instructopen2.9 GB · 3 benchmarks—–––9–32.9 GB
139Llama 3.1 8B Instructopen5.3 GB · 11 benchmarks14–14716115.3 GB

The range after each score is a 90% interval: given the boards a model has results on, its true score is very likely inside it. Short ranges mean many agreeing results; long ranges mean few or conflicting ones.

Why some models are missing. A model gets a score only with results on at least 4 benchmarks across at least 2 categories, one of them reasoning or coding. Skill scores need at least 2 benchmarks in that skill; otherwise they show as –.

VRAM is llmrun's estimate at Q4_K_M (or the smallest quantization we track) for the weights plus a working context. Proprietary models can't run locally, so the VRAM filter hides them.

How the llmrun Score is built · llmrun does not run these benchmarks; scores are aggregated from public sources.