Best open LLMs for agents
Open models ranked by agentic tool use, from agentic and tool-use boards such as Terminal-Bench and tau-bench. The score shown is the overall llmrun Score; the highlighted column is the agentic tool use skill score used for the order.
40 benchmarks4 sourcesUpdated 3 Oct 2026Reference scale 2026-Q4How the score is built
| # | Model | llmrun Score | Coding | Agents | Math | Science | Reasoning | Benchmarks | VRAM |
|---|---|---|---|---|---|---|---|---|---|
| 1 | GLM 5.3456 GB · 21 benchmarks | 67 | 68 | 72 | 49 | 58 | 21 | 456 GB | |
| 2 | GLM 5.3 Flash195 GB · 17 benchmarks | 52 | 63 | 62 | 44 | 37 | 17 | 195 GB | |
| 3 | Kimi K31674 GB · 26 benchmarks | 73 | 61 | 81 | 52 | 63 | 26 | 1674 GB | |
| 4 | GLM 5.2456 GB · 24 benchmarks | 58 | 60 | 68 | 47 | 52 | 24 | 456 GB | |
| 5 | GLM 5.1457 GB · 13 benchmarks | 53 | 54 | 57 | 40 | 46 | 13 | 457 GB | |
| 6 | Kimi K2.7 Code620 GB · 19 benchmarks | 53 | 46 | 61 | 40 | 50 | 19 | 620 GB | |
| 7 | GLM 5457 GB · 15 benchmarks | 49 | 45 | – | – | 40 | 15 | 457 GB | |
| 8 | MiniMax M3257 GB · 19 benchmarks | 52 | 40 | 46 | 41 | 38 | 19 | 257 GB | |
| 9 | Kimi K2.5620 GB · 20 benchmarks | 49 | 37 | – | 40 | 45 | 20 | 620 GB | |
| 10 | MiniMax M2.5138 GB · 11 benchmarks | 47 | 37 | – | – | 45 | 11 | 138 GB | |
| 11 | DeepSeek V3.2415 GB · 13 benchmarks | 35 | 36 | – | – | 43 | 13 | 415 GB | |
| 12 | Kimi K2 Thinking620 GB · 9 benchmarks | 43 | 34 | – | – | – | 9 | 620 GB | |
| 13 | GLM 4.7216 GB · 13 benchmarks | 41 | 32 | 51 | 36 | 35 | 13 | 216 GB | |
| 14 | GLM 4.5216 GB · 7 benchmarks | 38 | 32 | – | – | – | 7 | 216 GB | |
| 15 | Qwen3 Coder 480B A35B Instruct289 GB · 6 benchmarks | — | 40 | 31 | – | – | – | 6 | 289 GB |
| 16 | MiniMax M2138 GB · 5 benchmarks | – | 31 | – | – | – | 5 | 138 GB | |
| 17 | GLM 4.6215 GB · 6 benchmarks | 36 | 31 | – | 29 | – | 6 | 215 GB | |
| 18 | Kimi K2 Instruct620 GB · 10 benchmarks | 41 | 28 | – | – | 25 | 10 | 620 GB | |
| 19 | Qwen2.5 Coder 32B Instruct20.5 GB · 3 benchmarks | — | – | 15 | – | – | – | 3 | 20.5 GB |
| 20 | GPT OSS 120B70.5 GB · 20 benchmarks | 37 | 14 | – | 31 | 30 | 20 | 70.5 GB |
The range after each score is a 90% interval: given the boards a model has results on, its true score is very likely inside it. Short ranges mean many agreeing results; long ranges mean few or conflicting ones.
Why some models are missing. A model gets a score only with results on at least 4 benchmarks across at least 2 categories, one of them reasoning or coding. Skill scores need at least 2 benchmarks in that skill; otherwise they show as –.
VRAM is llmrun's estimate at Q4_K_M (or the smallest quantization we track) for the weights plus a working context. Proprietary models can't run locally, so the VRAM filter hides them.
How the llmrun Score is built · llmrun does not run these benchmarks; scores are aggregated from public sources.