Devstral Small 2 24B Instruct 2512 — Benchmarks

Benchmark scores for Devstral Small 2 24B Instruct 2512 aggregated from public leaderboards, with how it ranks among open models. See hardware requirements for what you need to run it.

No overall rank for Devstral Small 2 24B Instruct 2512: an overall score is only published when a model has been measured on enough of the benchmarks current models are still submitted to. Its individual scores below stand on their own — see the methodology for how the overall score is built.

Coding

BenchmarkScoreOpen rankAll models
SciCode28.8%#50 / 62#145 / 161
SWE-bench Verified56.4%#14 / 20#84 / 162

Reasoning

BenchmarkScoreOpen rankAll models
CritPt0.0%#54 / 62#160 / 172

Scores aggregated from public benchmark sources (each linked from the benchmark pages). llmrun does not run these benchmarks.