DeepSeek V3.1 — Benchmarks
Benchmark scores for DeepSeek V3.1 aggregated from public leaderboards, with how it ranks among open models. See hardware requirements for what you need to run it.
Overall rank: #17 of 78 open modelscomposite 55.5/100 across 5 benchmarks in 3 categories · methodology
Coding
| Benchmark | Score | Open rank | All models |
|---|---|---|---|
| WeirdML | 38.4% | #27 / 46 | #115 / 161 |
Instruction Following
| Benchmark | Score | Open rank | All models |
|---|---|---|---|
| LMArena Text | 1417.3 | #34 / 182 | #118 / 391 |
Knowledge
| Benchmark | Score | Open rank | All models |
|---|---|---|---|
| MMLU-Pro | 84.8% | #9 / 144 | #35 / 257 |
Reasoning
| Benchmark | Score | Open rank | All models |
|---|---|---|---|
| DTBench | 82.7% | #18 / 67 | #97 / 210 |
| LMCA | 24.3% | #29 / 53 | #129 / 172 |
| SimpleBench | 40.0% | #14 / 25 | #72 / 101 |
| Fiction.LiveBench | 52.8% | #16 / 24 | #40 / 58 |
Scores aggregated from public benchmark sources (each linked from the benchmark pages). llmrun does not run these benchmarks.