DeepSeek v3 — Benchmarks

Benchmark scores for DeepSeek v3 aggregated from public leaderboards, with how it ranks among open models. See hardware requirements for what you need to run it.

Overall rank: #33 of 73 open modelscomposite 50/100 across 11 benchmarks in 4 categories · methodology

Coding

BenchmarkScoreOpen rankAll models
Aider Polyglot48.4%#9 / 18#40 / 69
SWE-bench Lite36.7%#2 / 3#34 / 80

Knowledge

BenchmarkScoreOpen rankAll models
MMLU87.2%#1 / 76#3 / 136
HellaSwag88.9%#3 / 42#5 / 76
MMLU-Pro75.9%#27 / 119#83 / 259

Math

BenchmarkScoreOpen rankAll models
AIME 2024/202515.8%#18 / 34#113 / 155
MATH Level 564.8%#10 / 32#50 / 108
FrontierMath1.7%#10 / 12#85 / 101

Reasoning

BenchmarkScoreOpen rankAll models
GPQA Diamond56.5%#16 / 46#109 / 182
BIG-Bench Hard87.5%#1 / 37#2 / 50
SimpleBench18.9%#19 / 19#85 / 90

Scores aggregated from public benchmark sources (each linked from the benchmark pages). llmrun does not run these benchmarks.

DeepSeek v3 Benchmarks — Scores & Rankings | llmrun