Kimi K2.5 — Benchmarks
Benchmark scores for Kimi K2.5 aggregated from public leaderboards, with how it ranks among open models. See hardware requirements for what you need to run it.
Overall rank: #13 of 78 open modelscomposite 58/100 across 15 benchmarks in 4 categories · methodology
Coding
| Benchmark | Score | Open rank | All models |
|---|---|---|---|
| SciCode | 49.0% | #11 / 62 | #72 / 161 |
| WeirdML | 45.6% | #15 / 46 | #85 / 161 |
| SWE-bench Multilingual | 67.3% | #3 / 4 | #7 / 13 |
| LMArena WebDev | 1436.3 | #19 / 48 | #64 / 123 |
| Terminal-Bench | 43.2% | #3 / 16 | #30 / 57 |
| SWE-bench Verified | 70.8% | #4 / 20 | #37 / 162 |
Instruction Following
| Benchmark | Score | Open rank | All models |
|---|---|---|---|
| LMArena Text | 1450.4 | #13 / 182 | #68 / 391 |
Knowledge
| Benchmark | Score | Open rank | All models |
|---|---|---|---|
| Humanity's Last Exam | 24.4% | #1 / 4 | #16 / 49 |
| SimpleQA | 34.3% | #9 / 15 | #47 / 80 |
| MMLU-Pro | 87.1% | #3 / 144 | #17 / 257 |
Math
| Benchmark | Score | Open rank | All models |
|---|---|---|---|
| AIME 2024/2025 | 92.2% | #9 / 82 | #54 / 270 |
Reasoning
| Benchmark | Score | Open rank | All models |
|---|---|---|---|
| Chess Puzzles | 12.0% | #20 / 65 | #124 / 205 |
| GPQA Diamond | 87.6% | #15 / 94 | #61 / 289 |
| SimpleBench | 46.8% | #10 / 25 | #57 / 101 |
| CritPt | 3.1% | #20 / 62 | #95 / 172 |
| ARC-AGI | 65.3% | #7 / 16 | #92 / 200 |
| ARC-AGI-2 | 11.8% | #7 / 16 | #92 / 190 |
| Fiction.LiveBench | 86.1% | #1 / 24 | #7 / 58 |
| CL-bench | 19.3% | #1 / 7 | #10 / 22 |
| CL-bench Life | 13.2% | #2 / 6 | #10 / 17 |
Scores aggregated from public benchmark sources (each linked from the benchmark pages). llmrun does not run these benchmarks.