Best open LLMs for agents

Open models ranked by agentic tool use, from agentic and tool-use boards such as Terminal-Bench and tau-bench. The score shown is the overall llmrun Score; the highlighted column is the agentic tool use skill score used for the order.

40 benchmarks4 sourcesUpdated 3 Oct 2026Reference scale 2026-Q4How the score is built

Models ranked by agents llmrun Score
#Modelllmrun ScoreCodingAgentsMathScienceReasoningBenchmarksVRAM
1GLM 5.3456 GB · 21 benchmarks676872495821456 GB
2GLM 5.3 Flash195 GB · 17 benchmarks526362443717195 GB
3Kimi K31674 GB · 26 benchmarks7361815263261674 GB
4GLM 5.2456 GB · 24 benchmarks586068475224456 GB
5GLM 5.1457 GB · 13 benchmarks535457404613457 GB
6Kimi K2.7 Code620 GB · 19 benchmarks534661405019620 GB
7GLM 5457 GB · 15 benchmarks4945––4015457 GB
8MiniMax M3257 GB · 19 benchmarks524046413819257 GB
9Kimi K2.5620 GB · 20 benchmarks4937–404520620 GB
10MiniMax M2.5138 GB · 11 benchmarks4737––4511138 GB
11DeepSeek V3.2415 GB · 13 benchmarks3536––4313415 GB
12Kimi K2 Thinking620 GB · 9 benchmarks4334–––9620 GB
13GLM 4.7216 GB · 13 benchmarks413251363513216 GB
14GLM 4.5216 GB · 7 benchmarks3832–––7216 GB
15Qwen3 Coder 480B A35B Instruct289 GB · 6 benchmarks—4031–––6289 GB
16MiniMax M2138 GB · 5 benchmarks–31–––5138 GB
17GLM 4.6215 GB · 6 benchmarks3631–29–6215 GB
18Kimi K2 Instruct620 GB · 10 benchmarks4128––2510620 GB
19Qwen2.5 Coder 32B Instruct20.5 GB · 3 benchmarks—–15–––320.5 GB
20GPT OSS 120B70.5 GB · 20 benchmarks3714–31302070.5 GB

The range after each score is a 90% interval: given the boards a model has results on, its true score is very likely inside it. Short ranges mean many agreeing results; long ranges mean few or conflicting ones.

Why some models are missing. A model gets a score only with results on at least 4 benchmarks across at least 2 categories, one of them reasoning or coding. Skill scores need at least 2 benchmarks in that skill; otherwise they show as –.

VRAM is llmrun's estimate at Q4_K_M (or the smallest quantization we track) for the weights plus a working context. Proprietary models can't run locally, so the VRAM filter hides them.

How the llmrun Score is built · llmrun does not run these benchmarks; scores are aggregated from public sources.