Kimi K2.5 — Benchmarks

Benchmark scores for Kimi K2.5 aggregated from public leaderboards, with how it ranks among open models. See hardware requirements for what you need to run it.

Overall rank: #13 of 78 open modelscomposite 58/100 across 15 benchmarks in 4 categories · methodology

Coding

BenchmarkScoreOpen rankAll models
SciCode49.0%#11 / 62#72 / 161
WeirdML45.6%#15 / 46#85 / 161
SWE-bench Multilingual67.3%#3 / 4#7 / 13
LMArena WebDev1436.3#19 / 48#64 / 123
Terminal-Bench43.2%#3 / 16#30 / 57
SWE-bench Verified70.8%#4 / 20#37 / 162

Instruction Following

BenchmarkScoreOpen rankAll models
LMArena Text1450.4#13 / 182#68 / 391

Knowledge

BenchmarkScoreOpen rankAll models
Humanity's Last Exam24.4%#1 / 4#16 / 49
SimpleQA34.3%#9 / 15#47 / 80
MMLU-Pro87.1%#3 / 144#17 / 257

Math

BenchmarkScoreOpen rankAll models
AIME 2024/202592.2%#9 / 82#54 / 270

Reasoning

BenchmarkScoreOpen rankAll models
Chess Puzzles12.0%#20 / 65#124 / 205
GPQA Diamond87.6%#15 / 94#61 / 289
SimpleBench46.8%#10 / 25#57 / 101
CritPt3.1%#20 / 62#95 / 172
ARC-AGI65.3%#7 / 16#92 / 200
ARC-AGI-211.8%#7 / 16#92 / 190
Fiction.LiveBench86.1%#1 / 24#7 / 58
CL-bench19.3%#1 / 7#10 / 22
CL-bench Life13.2%#2 / 6#10 / 17

Scores aggregated from public benchmark sources (each linked from the benchmark pages). llmrun does not run these benchmarks.