GPT OSS 120B — Benchmarks

Benchmark scores for GPT OSS 120B aggregated from public leaderboards, with how it ranks among open models. See hardware requirements for what you need to run it.

Overall rank: #27 of 73 open modelscomposite 52/100 across 9 benchmarks in 4 categories · methodology

Coding

BenchmarkScoreOpen rankAll models
Terminal-Bench18.7%#14 / 16#52 / 57
Aider Polyglot41.8%#10 / 18#45 / 69
SWE-bench Bash Only26.0%#8 / 9#42 / 48
SWE-bench Verified26.0%#13 / 13#143 / 163

Knowledge

BenchmarkScoreOpen rankAll models
SimpleQA13.9%#10 / 11#59 / 65
MMLU-Pro80.8%#21 / 119#61 / 259

Math

BenchmarkScoreOpen rankAll models
AIME 2024/202588.9%#6 / 34#34 / 155

Reasoning

BenchmarkScoreOpen rankAll models
GPQA Diamond75.8%#11 / 46#77 / 182
SimpleBench22.1%#17 / 19#83 / 90

Scores aggregated from public benchmark sources (each linked from the benchmark pages). llmrun does not run these benchmarks.