DeepSeek v3 — Benchmarks
Benchmark scores for DeepSeek v3 aggregated from public leaderboards, with how it ranks among open models. See hardware requirements for what you need to run it.
Overall rank: #41 of 78 open modelscomposite 43.5/100 across 9 benchmarks in 4 categories · methodology
Coding
| Benchmark | Score | Open rank | All models |
|---|---|---|---|
| SciCode | 35.4% | #41 / 62 | #135 / 161 |
| Aider Polyglot | 48.4% | #9 / 18 | #40 / 69 |
| SWE-bench Lite | 36.7% | #2 / 3 | #34 / 80 |
Instruction Following
| Benchmark | Score | Open rank | All models |
|---|---|---|---|
| LMArena Text | 1358.4 | #65 / 182 | #185 / 391 |
Knowledge
| Benchmark | Score | Open rank | All models |
|---|---|---|---|
| MMLU-Pro | 75.9% | #29 / 144 | #82 / 257 |
| ARC Challenge | 95.3% | #1 / 56 | #1 / 76 |
| TriviaQA | 82.9% | #2 / 19 | #10 / 39 |
| MMLU | 87.2% | #1 / 85 | #3 / 136 |
| HellaSwag | 88.9% | #3 / 47 | #5 / 75 |
Math
| Benchmark | Score | Open rank | All models |
|---|---|---|---|
| MATH Level 5 | 64.8% | #10 / 39 | #50 / 107 |
| AIME 2024/2025 | 15.8% | #54 / 82 | #217 / 270 |
Reasoning
| Benchmark | Score | Open rank | All models |
|---|---|---|---|
| GPQA Diamond | 56.5% | #43 / 94 | #199 / 289 |
| SimpleBench | 18.9% | #24 / 25 | #96 / 101 |
| CritPt | 0.0% | #43 / 62 | #138 / 172 |
| WinoGrande | 85.2% | #4 / 51 | #7 / 79 |
| BIG-Bench Hard | 87.5% | #1 / 38 | #2 / 50 |
Scores aggregated from public benchmark sources (each linked from the benchmark pages). llmrun does not run these benchmarks.