Best LLMs: open and proprietary
Open and proprietary models together on one 0–100 scale (overall capability). Switch to local models to see only what you can download and run.
40 benchmarks4 sourcesUpdated 3 Oct 2026Reference scale 2026-Q4How the score is built
| # | Model | llmrun Score | Coding | Agents | Math | Science | Reasoning | Benchmarks | VRAM |
|---|---|---|---|---|---|---|---|---|---|
| 241 | Mixtral 8x7B Instruct v0.1open28.6 GB · 5 benchmarks | – | – | – | – | 16 | 5 | 28.6 GB | |
| 242 | Gemini 1.0 Pro 0014 benchmarks | – | – | 11 | – | – | 4 | – | |
| 243 | Meta Llama 3 8B Instructopen5.3 GB · 7 benchmarks | – | – | 10 | – | 14 | 7 | 5.3 GB | |
| 244 | Llama 3.2 1B Instructopen0.8 GB · 4 benchmarks | – | – | – | – | – | 4 | 0.8 GB | |
| 245 | Deepseek Llm 67B Chatopen41.3 GB · 4 benchmarks | – | – | 10 | – | – | 4 | 41.3 GB | |
| 246 | GPT 3.5 Turbo (Jan 25)10 benchmarks | – | – | 12 | – | 12 | 10 | – | |
| 247 | Mistral 7B Instruct v0.3open4.9 GB · 4 benchmarks | – | – | 10 | – | – | 4 | 4.9 GB | |
| 248 | Llama 2 70B Chat HFopen45.5 GB · 5 benchmarks | – | – | 9 | – | 10 | 5 | 45.5 GB | |
| 249 | Gemma 3 1B ITopen0.7 GB · 4 benchmarks | – | – | – | – | – | 4 | 0.7 GB |
The range after each score is a 90% interval: given the boards a model has results on, its true score is very likely inside it. Short ranges mean many agreeing results; long ranges mean few or conflicting ones.
Why some models are missing. A model gets a score only with results on at least 4 benchmarks across at least 2 categories, one of them reasoning or coding. Skill scores need at least 2 benchmarks in that skill; otherwise they show as –.
VRAM is llmrun's estimate at Q4_K_M (or the smallest quantization we track) for the weights plus a working context. Proprietary models can't run locally, so the VRAM filter hides them.
How the llmrun Score is built · llmrun does not run these benchmarks; scores are aggregated from public sources.