Kimi K2 Thinking — Benchmarks

Benchmark scores for Kimi K2 Thinking aggregated from public leaderboards, with how it ranks among open models. See hardware requirements for what you need to run it.

Overall rank: #24 of 73 open modelscomposite 54.2/100 across 7 benchmarks in 4 categories · methodology

Coding

BenchmarkScoreOpen rankAll models
SWE-bench Verified63.4%#6 / 13#67 / 163
SWE-bench Bash Only63.4%#1 / 9#24 / 48
Terminal-Bench35.7%#7 / 16#37 / 57

Knowledge

BenchmarkScoreOpen rankAll models
SimpleQA31.6%#7 / 11#45 / 65

Math

BenchmarkScoreOpen rankAll models
AIME 2024/202583.1%#9 / 34#53 / 155
FrontierMath21.4%#5 / 12#37 / 101

Reasoning

BenchmarkScoreOpen rankAll models
GPQA Diamond84.2%#7 / 46#46 / 182

Scores aggregated from public benchmark sources (each linked from the benchmark pages). llmrun does not run these benchmarks.

Kimi K2 Thinking Benchmarks — Scores & Rankings | llmrun