DeepSeek R1 — Benchmarks

Benchmark scores for DeepSeek R1 aggregated from public leaderboards, with how it ranks among open models. See hardware requirements for what you need to run it.

Overall rank: #16 of 73 open modelscomposite 61.9/100 across 7 benchmarks in 4 categories · methodology

Coding

BenchmarkScoreOpen rankAll models
Aider Polyglot56.9%#7 / 18#31 / 69

Knowledge

BenchmarkScoreOpen rankAll models
MMLU-Pro84.0%#10 / 119#41 / 259

Math

BenchmarkScoreOpen rankAll models
AIME 2024/202553.3%#12 / 34#86 / 155
MATH Level 593.0%#2 / 32#19 / 108

Reasoning

BenchmarkScoreOpen rankAll models
SimpleBench30.9%#12 / 19#69 / 90
GPQA Diamond69.2%#13 / 46#88 / 182
ARC-AGI15.8%#6 / 10#134 / 158

Scores aggregated from public benchmark sources (each linked from the benchmark pages). llmrun does not run these benchmarks.