DeepSeek R1 — Benchmarks

Benchmark scores for DeepSeek R1 aggregated from public leaderboards, with how it ranks among open models. See hardware requirements for what you need to run it.

Overall rank: #34 of 123 open modelscomposite 57.1/100 across 9 benchmarks in 4 categories · methodology

Coding

BenchmarkScoreOpen rankAll models
Aider Polyglot56.9%#7 / 18#31 / 69
SciCode35.7%#32 / 50#127 / 155

Knowledge

BenchmarkScoreOpen rankAll models
MMLU-Pro84.0%#10 / 119#41 / 259

Math

BenchmarkScoreOpen rankAll models
AIME 2024/202553.3%#39 / 74#167 / 270
MATH Level 593.0%#2 / 32#19 / 108

Reasoning

BenchmarkScoreOpen rankAll models
GPQA Diamond69.2%#30 / 83#164 / 291
ARC-AGI15.8%#12 / 16#174 / 200
SimpleBench30.9%#16 / 23#80 / 101
CritPt1.1%#27 / 50#110 / 167

Scores aggregated from public benchmark sources (each linked from the benchmark pages). llmrun does not run these benchmarks.