DeepSeek v3 — Benchmarks

Benchmark scores for DeepSeek v3 aggregated from public leaderboards, with how it ranks among open models. See hardware requirements for what you need to run it.

Overall rank: #41 of 78 open modelscomposite 43.5/100 across 9 benchmarks in 4 categories · methodology

Coding

BenchmarkScoreOpen rankAll models
SciCode35.4%#41 / 62#135 / 161
Aider Polyglot48.4%#9 / 18#40 / 69
SWE-bench Lite36.7%#2 / 3#34 / 80

Instruction Following

BenchmarkScoreOpen rankAll models
LMArena Text1358.4#65 / 182#185 / 391

Knowledge

BenchmarkScoreOpen rankAll models
MMLU-Pro75.9%#29 / 144#82 / 257
ARC Challenge95.3%#1 / 56#1 / 76
TriviaQA82.9%#2 / 19#10 / 39
MMLU87.2%#1 / 85#3 / 136
HellaSwag88.9%#3 / 47#5 / 75

Math

BenchmarkScoreOpen rankAll models
MATH Level 564.8%#10 / 39#50 / 107
AIME 2024/202515.8%#54 / 82#217 / 270

Reasoning

BenchmarkScoreOpen rankAll models
GPQA Diamond56.5%#43 / 94#199 / 289
SimpleBench18.9%#24 / 25#96 / 101
CritPt0.0%#43 / 62#138 / 172
WinoGrande85.2%#4 / 51#7 / 79
BIG-Bench Hard87.5%#1 / 38#2 / 50

Scores aggregated from public benchmark sources (each linked from the benchmark pages). llmrun does not run these benchmarks.