DeepSeek V3.2 — Benchmarks
Benchmark scores for DeepSeek V3.2 aggregated from public leaderboards, with how it ranks among open models. See hardware requirements for what you need to run it.
Overall rank: #50 of 123 open modelscomposite 51.8/100 across 9 benchmarks in 4 categories · methodology
Coding
| Benchmark | Score | Open rank | All models |
|---|---|---|---|
| SWE-bench Lite | 30.7% | #3 / 3 | #46 / 80 |
| Terminal-Bench | 39.6% | #5 / 16 | #33 / 57 |
| SWE-bench Verified | 70.0% | #5 / 16 | #44 / 162 |
| SWE-bench Bash Only | 60.0% | #3 / 9 | #26 / 48 |
| SWE-bench Multilingual | 59.0% | #4 / 4 | #12 / 13 |
Knowledge
| Benchmark | Score | Open rank | All models |
|---|---|---|---|
| MMLU-Pro | 85.0% | #6 / 119 | #33 / 259 |
Math
| Benchmark | Score | Open rank | All models |
|---|---|---|---|
| FrontierMath | 22.1% | #4 / 12 | #36 / 101 |
| ProofBench | 8.0% | #14 / 21 | #54 / 64 |
Reasoning
| Benchmark | Score | Open rank | All models |
|---|---|---|---|
| ARC-AGI | 57.0% | #9 / 16 | #103 / 200 |
Scores aggregated from public benchmark sources (each linked from the benchmark pages). llmrun does not run these benchmarks.