Kimi K2 Thinking — Benchmarks

Benchmark scores for Kimi K2 Thinking aggregated from public leaderboards, with how it ranks among open models. See hardware requirements for what you need to run it.

Overall rank: #24 of 123 open modelscomposite 61.8/100 across 6 benchmarks in 3 categories · methodology

Coding

BenchmarkScoreOpen rankAll models
Terminal-Bench35.7%#7 / 16#37 / 57
SWE-bench Verified63.4%#10 / 16#67 / 162
SWE-bench Bash Only63.4%#1 / 9#24 / 48

Math

BenchmarkScoreOpen rankAll models
AIME 2024/202583.1%#20 / 74#92 / 270
FrontierMath21.4%#5 / 12#37 / 101

Reasoning

BenchmarkScoreOpen rankAll models
GPQA Diamond84.2%#19 / 83#84 / 291

Scores aggregated from public benchmark sources (each linked from the benchmark pages). llmrun does not run these benchmarks.