Kimi K2 Thinking — Benchmarks
Benchmark scores for Kimi K2 Thinking aggregated from public leaderboards, with how it ranks among open models. See hardware requirements for what you need to run it.
Overall rank: #24 of 73 open modelscomposite 54.2/100 across 7 benchmarks in 4 categories · methodology
Coding
| Benchmark | Score | Open rank | All models |
|---|---|---|---|
| SWE-bench Verified | 63.4% | #6 / 13 | #67 / 163 |
| SWE-bench Bash Only | 63.4% | #1 / 9 | #24 / 48 |
| Terminal-Bench | 35.7% | #7 / 16 | #37 / 57 |
Knowledge
| Benchmark | Score | Open rank | All models |
|---|---|---|---|
| SimpleQA | 31.6% | #7 / 11 | #45 / 65 |
Math
| Benchmark | Score | Open rank | All models |
|---|---|---|---|
| AIME 2024/2025 | 83.1% | #9 / 34 | #53 / 155 |
| FrontierMath | 21.4% | #5 / 12 | #37 / 101 |
Reasoning
| Benchmark | Score | Open rank | All models |
|---|---|---|---|
| GPQA Diamond | 84.2% | #7 / 46 | #46 / 182 |
Scores aggregated from public benchmark sources (each linked from the benchmark pages). llmrun does not run these benchmarks.