DeepSeek V4 Flash 0731 — Benchmarks
Benchmark scores for DeepSeek V4 Flash 0731 aggregated from public leaderboards, with how it ranks among open models. See hardware requirements for what you need to run it.
Overall rank: #25 of 123 open modelscomposite 60.9/100 across 11 benchmarks in 4 categories · methodology
Coding
| Benchmark | Score | Open rank | All models |
|---|---|---|---|
| LiveBench Coding | 75.0 | #8 / 16 | #37 / 54 |
| SciCode | 49.9% | #7 / 50 | #60 / 155 |
Knowledge
| Benchmark | Score | Open rank | All models |
|---|---|---|---|
| SimpleQA | 33.6% | #12 / 15 | #49 / 78 |
Math
| Benchmark | Score | Open rank | All models |
|---|---|---|---|
| LiveBench Math | 86.8 | #7 / 16 | #36 / 54 |
| AIME 2024/2025 | 94.4% | #6 / 74 | #39 / 270 |
| ProofBench | 56.0% | #2 / 21 | #13 / 64 |
Reasoning
| Benchmark | Score | Open rank | All models |
|---|---|---|---|
| LiveBench Reasoning | 86.6 | #3 / 16 | #24 / 54 |
| GPQA Diamond | 91.0% | #4 / 83 | #26 / 291 |
| ARC-AGI | 89.0% | #3 / 16 | #51 / 200 |
| SimpleBench | 61.1% | #1 / 23 | #27 / 101 |
| CritPt | 16.6% | #5 / 50 | #48 / 167 |
Scores aggregated from public benchmark sources (each linked from the benchmark pages). llmrun does not run these benchmarks.