Devstral Small 2 24B Instruct 2512 — Benchmarks
Benchmark scores for Devstral Small 2 24B Instruct 2512 aggregated from public leaderboards, with how it ranks among open models. See hardware requirements for what you need to run it.
No overall rank for Devstral Small 2 24B Instruct 2512: an overall score is only published when a model has been measured on enough of the benchmarks current models are still submitted to. Its individual scores below stand on their own — see the methodology for how the overall score is built.
Coding
| Benchmark | Score | Open rank | All models |
|---|---|---|---|
| SciCode | 28.8% | #50 / 62 | #145 / 161 |
| SWE-bench Verified | 56.4% | #14 / 20 | #84 / 162 |
Reasoning
| Benchmark | Score | Open rank | All models |
|---|---|---|---|
| CritPt | 0.0% | #54 / 62 | #160 / 172 |
Scores aggregated from public benchmark sources (each linked from the benchmark pages). llmrun does not run these benchmarks.