DeepSeek V3.2 — Benchmarks

Benchmark scores for DeepSeek V3.2 aggregated from public leaderboards, with how it ranks among open models. See hardware requirements for what you need to run it.

Overall rank: #50 of 123 open modelscomposite 51.8/100 across 9 benchmarks in 4 categories · methodology

Coding

BenchmarkScoreOpen rankAll models
SWE-bench Lite30.7%#3 / 3#46 / 80
Terminal-Bench39.6%#5 / 16#33 / 57
SWE-bench Verified70.0%#5 / 16#44 / 162
SWE-bench Bash Only60.0%#3 / 9#26 / 48
SWE-bench Multilingual59.0%#4 / 4#12 / 13

Knowledge

BenchmarkScoreOpen rankAll models
MMLU-Pro85.0%#6 / 119#33 / 259

Math

BenchmarkScoreOpen rankAll models
FrontierMath22.1%#4 / 12#36 / 101
ProofBench8.0%#14 / 21#54 / 64

Reasoning

BenchmarkScoreOpen rankAll models
ARC-AGI57.0%#9 / 16#103 / 200

Scores aggregated from public benchmark sources (each linked from the benchmark pages). llmrun does not run these benchmarks.