Best LLMs for agents: open and proprietary
Open and proprietary models together on one 0–100 scale (agentic tool use). Switch to local models to see only what you can download and run.
40 benchmarks4 sourcesUpdated 3 Oct 2026Reference scale 2026-Q4How the score is built
| # | Model | llmrun Score | Coding | Agents | Math | Science | Reasoning | Benchmarks | VRAM |
|---|---|---|---|---|---|---|---|---|---|
| 1 | GPT 6 Astra6 benchmarks | – | 87 | – | – | 74 | 6 | – | |
| 2 | Claude Opus 4.77 benchmarks | 68 | 75 | – | – | 54 | 7 | – | |
| 3 | Claude Opus 5 Max22 benchmarks | 78 | 72 | 91 | 55 | 75 | 22 | – | |
| 4 | GPT 6 Astra Max19 benchmarks | 83 | 72 | 96 | 57 | 86 | 19 | – | |
| 5 | Grok 4.65 benchmarks | – | 72 | – | – | – | 5 | – | |
| 6 | GPT 5.6 Sol Max21 benchmarks | 78 | 70 | 91 | 57 | 77 | 21 | – | |
| 7 | GLM 5.3open456 GB · 21 benchmarks | 67 | 68 | 72 | 49 | 58 | 21 | 456 GB | |
| 8 | GPT 5.6 Terra Max20 benchmarks | 73 | 67 | 87 | 55 | 71 | 20 | – | |
| 9 | Claude Opus 4.68 benchmarks | 64 | 64 | – | – | – | 8 | – | |
| 10 | Claude Opus 4.85 benchmarks | – | 64 | – | – | 56 | 5 | – | |
| 11 | GLM 5.3 Flashopen195 GB · 17 benchmarks | 52 | 63 | 62 | 44 | 37 | 17 | 195 GB | |
| 12 | GPT 5.6 Luna Max20 benchmarks | 69 | 63 | 82 | 48 | 62 | 20 | – | |
| 13 | Claude Sonnet 55 benchmarks | – | 62 | – | – | – | 5 | – | |
| 14 | Claude Fable 5 Max20 benchmarks | 78 | 62 | 94 | 54 | 74 | 20 | – | |
| 15 | Kimi K3open1674 GB · 26 benchmarks | 73 | 61 | 81 | 52 | 63 | 26 | 1674 GB | |
| 16 | GLM 5.2open456 GB · 24 benchmarks | 58 | 60 | 68 | 47 | 52 | 24 | 456 GB | |
| 17 | Muse Spark 1.19 benchmarks | 66 | 59 | 69 | 48 | – | 9 | – | |
| 18 | GPT 5.4 (Mar 05, 2026)20 benchmarks | 69 | 57 | 78 | 51 | 67 | 20 | – | |
| 19 | Gemini 3 Pro Preview17 benchmarks | 62 | 55 | 60 | 46 | 54 | 17 | – | |
| 20 | Claude Opus 4.8 Max19 benchmarks | 64 | 55 | 82 | 48 | 66 | 19 | – | |
| 21 | GLM 5.1open457 GB · 13 benchmarks | 53 | 54 | 57 | 40 | 46 | 13 | 457 GB | |
| 22 | Claude Opus 4.5 (Nov 01, 2025)9 benchmarks | – | 52 | – | – | 55 | 9 | – | |
| 23 | Grok 4.2011 benchmarks | 55 | 49 | – | – | 63 | 11 | – | |
| 24 | GPT 5.2 Codex8 benchmarks | 62 | 48 | – | – | – | 8 | – | |
| 25 | Kimi K2.7 Codeopen620 GB · 19 benchmarks | 53 | 46 | 61 | 40 | 50 | 19 | 620 GB | |
| 26 | GLM 5open457 GB · 15 benchmarks | 49 | 45 | – | – | 40 | 15 | 457 GB | |
| 27 | Gemini 3 Pro Preview (2025-11-18)2 benchmarks | — | – | 44 | – | – | – | 2 | – |
| 28 | Claude Sonnet 4.68 benchmarks | – | 43 | – | – | 60 | 8 | – | |
| 29 | Claude Sonnet 4.5 (Sep 29, 2025)11 benchmarks | – | 42 | – | 32 | 46 | 11 | – | |
| 30 | Claude 4.5 Sonnet (20250929)2 benchmarks | — | – | 41 | – | – | – | 2 | – |
| 31 | Gemini 3 Pro3 benchmarks | — | – | 40 | – | – | – | 3 | – |
| 32 | MiniMax M3open257 GB · 19 benchmarks | 52 | 40 | 46 | 41 | 38 | 19 | 257 GB | |
| 33 | Claude 4 Opus (20250514)2 benchmarks | — | – | 38 | – | – | – | 2 | – |
| 34 | Kimi K2.5open620 GB · 20 benchmarks | 49 | 37 | – | 40 | 45 | 20 | 620 GB | |
| 35 | MiniMax M2.5open138 GB · 11 benchmarks | 47 | 37 | – | – | 45 | 11 | 138 GB | |
| 36 | Claude 4 Sonnet (20250514)2 benchmarks | — | – | 36 | – | – | – | 2 | – |
| 37 | DeepSeek V3.2open415 GB · 13 benchmarks | 35 | 36 | – | – | 43 | 13 | 415 GB | |
| 38 | Kimi K2 Thinkingopen620 GB · 9 benchmarks | 43 | 34 | – | – | – | 9 | 620 GB | |
| 39 | GLM 4.7open216 GB · 13 benchmarks | 41 | 32 | 51 | 36 | 35 | 13 | 216 GB | |
| 40 | GLM 4.5open216 GB · 7 benchmarks | 38 | 32 | – | – | – | 7 | 216 GB | |
| 41 | o3 (2025-04-16)2 benchmarks | — | – | 32 | – | – | – | 2 | – |
| 42 | Qwen3 Coder 480B A35B Instructopen289 GB · 6 benchmarks | — | 40 | 31 | – | – | – | 6 | 289 GB |
| 43 | MiniMax M2open138 GB · 5 benchmarks | – | 31 | – | – | – | 5 | 138 GB | |
| 44 | GLM 4.6open215 GB · 6 benchmarks | 36 | 31 | – | 29 | – | 6 | 215 GB | |
| 45 | Grok 4 (Jul 09)12 benchmarks | 49 | 30 | – | – | 48 | 12 | – | |
| 46 | Gemini 2.5 Pro13 benchmarks | – | 29 | 49 | 37 | 43 | 13 | – | |
| 47 | Gemini 2.5 Pro (2025-05-06)2 benchmarks | — | – | 28 | – | – | – | 2 | – |
| 48 | Kimi K2 Instructopen620 GB · 10 benchmarks | 41 | 28 | – | – | 25 | 10 | 620 GB | |
| 49 | Claude 3.7 Sonnet (20250219)2 benchmarks | — | – | 28 | – | – | – | 2 | – |
| 50 | Claude 3.5 Sonnet (Oct 22, 2024)11 benchmarks | 38 | 26 | 23 | 21 | 29 | 11 | – | |
| 51 | o4-mini (2025-04-16)2 benchmarks | — | – | 23 | – | – | – | 2 | – |
| 52 | Gemini 2.5 Flash9 benchmarks | – | 22 | – | – | 34 | 9 | – | |
| 53 | Gemini 2.5 Flash (2025-04-17)2 benchmarks | — | – | 15 | – | – | – | 2 | – |
| 54 | Qwen2.5 Coder 32B Instructopen20.5 GB · 3 benchmarks | — | – | 15 | – | – | – | 3 | 20.5 GB |
| 55 | GPT OSS 120Bopen70.5 GB · 20 benchmarks | 37 | 14 | – | 31 | 30 | 20 | 70.5 GB | |
| 56 | Llama 4 Maverick Instruct2 benchmarks | — | – | 11 | – | – | – | 2 | – |
| 57 | Llama 4 Scout Instruct2 benchmarks | — | – | 7 | – | – | – | 2 | – |
The range after each score is a 90% interval: given the boards a model has results on, its true score is very likely inside it. Short ranges mean many agreeing results; long ranges mean few or conflicting ones.
Why some models are missing. A model gets a score only with results on at least 4 benchmarks across at least 2 categories, one of them reasoning or coding. Skill scores need at least 2 benchmarks in that skill; otherwise they show as –.
VRAM is llmrun's estimate at Q4_K_M (or the smallest quantization we track) for the weights plus a working context. Proprietary models can't run locally, so the VRAM filter hides them.
How the llmrun Score is built · llmrun does not run these benchmarks; scores are aggregated from public sources.