Best LLMs for science: open and proprietary
Open and proprietary models together on one 0–100 scale (scientific knowledge). Switch to local models to see only what you can download and run.
40 benchmarks4 sourcesUpdated 3 Oct 2026Reference scale 2026-Q4How the score is built
| # | Model | llmrun Score | Coding | Agents | Math | Science | Reasoning | Benchmarks | VRAM |
|---|---|---|---|---|---|---|---|---|---|
| 121 | Gemini 1.5 Pro 0029 benchmarks | – | – | 28 | 22 | 21 | 9 | – | |
| 122 | Claude 3.5 Sonnet (Oct 22, 2024)11 benchmarks | 38 | 26 | 23 | 21 | 29 | 11 | – | |
| 123 | Qwen3 8Bopen5.5 GB · 8 benchmarks | – | – | – | 21 | 21 | 8 | 5.5 GB | |
| 124 | Devstral Small 2 24B Instruct 2512open15.1 GB · 3 benchmarks | — | – | – | – | 20 | – | 3 | 15.1 GB |
| 125 | Llama 4 Scout 17B 16E Instructopen71.7 GB · 12 benchmarks | – | – | 24 | 18 | 20 | 12 | 71.7 GB | |
| 126 | Magistral Small 2509open15.1 GB · 6 benchmarks | – | – | – | 18 | 22 | 6 | 15.1 GB | |
| 127 | Mistral Small 3.2 24B Instruct 2506open15.1 GB · 6 benchmarks | – | – | – | 18 | 22 | 6 | 15.1 GB | |
| 128 | GPT 4.1 Nano (Apr 14, 2025)13 benchmarks | 21 | – | 30 | 17 | 17 | 13 | – | |
| 129 | MiMo v2 Flashopen186 GB · 3 benchmarks | — | 40 | – | – | 17 | – | 3 | 186 GB |
| 130 | Granite 4.1 30Bopen18.2 GB · 2 benchmarks | — | – | – | – | 17 | – | 2 | 18.2 GB |
| 131 | GPT 4o (Nov 20, 2024)10 benchmarks | 25 | – | 21 | 17 | 24 | 10 | – | |
| 132 | Mistral Small 3.1 24B Instruct 2503open15.1 GB · 7 benchmarks | – | – | 21 | 17 | 21 | 7 | 15.1 GB | |
| 133 | Llama 3.3 70B Instructopen46.6 GB · 13 benchmarks | 22 | – | 19 | 17 | 21 | 13 | 46.6 GB | |
| 134 | Gemma 3 27B ITopen18.1 GB · 11 benchmarks | 18 | – | 29 | 16 | 17 | 11 | 18.1 GB | |
| 135 | Solar Pro 32 benchmarks | — | – | – | – | 16 | – | 2 | – |
| 136 | Claude 3.5 Haiku (Oct 22, 2024)9 benchmarks | 29 | – | 20 | 13 | – | 9 | – | |
| 137 | Gemma 3 12B ITopen8.0 GB · 8 benchmarks | – | – | – | 12 | 15 | 8 | 8.0 GB | |
| 138 | Phi 4 Mini Instructopen2.9 GB · 3 benchmarks | — | – | – | – | 9 | – | 3 | 2.9 GB |
| 139 | Llama 3.1 8B Instructopen5.3 GB · 11 benchmarks | 14 | – | 14 | 7 | 16 | 11 | 5.3 GB |
The range after each score is a 90% interval: given the boards a model has results on, its true score is very likely inside it. Short ranges mean many agreeing results; long ranges mean few or conflicting ones.
Why some models are missing. A model gets a score only with results on at least 4 benchmarks across at least 2 categories, one of them reasoning or coding. Skill scores need at least 2 benchmarks in that skill; otherwise they show as –.
VRAM is llmrun's estimate at Q4_K_M (or the smallest quantization we track) for the weights plus a working context. Proprietary models can't run locally, so the VRAM filter hides them.
How the llmrun Score is built · llmrun does not run these benchmarks; scores are aggregated from public sources.