GPT OSS 120B — Benchmarks
Benchmark scores for GPT OSS 120B aggregated from public leaderboards, with how it ranks among open models. See hardware requirements for what you need to run it.
Overall rank: #36 of 123 open modelscomposite 57/100 across 10 benchmarks in 4 categories · methodology
Coding
| Benchmark | Score | Open rank | All models |
|---|---|---|---|
| Terminal-Bench | 18.7% | #14 / 16 | #52 / 57 |
| Aider Polyglot | 41.8% | #10 / 18 | #45 / 69 |
| SciCode | 38.9% | #26 / 50 | #116 / 155 |
| SWE-bench Verified | 26.0% | #16 / 16 | #142 / 162 |
| SWE-bench Bash Only | 26.0% | #8 / 9 | #42 / 48 |
Knowledge
| Benchmark | Score | Open rank | All models |
|---|---|---|---|
| MMLU-Pro | 80.8% | #21 / 119 | #61 / 259 |
Math
| Benchmark | Score | Open rank | All models |
|---|---|---|---|
| AIME 2024/2025 | 88.9% | #13 / 74 | #63 / 270 |
Reasoning
| Benchmark | Score | Open rank | All models |
|---|---|---|---|
| GPQA Diamond | 75.8% | #25 / 83 | #139 / 291 |
| SimpleBench | 22.1% | #21 / 23 | #94 / 101 |
| CritPt | 1.1% | #25 / 50 | #108 / 167 |
Scores aggregated from public benchmark sources (each linked from the benchmark pages). llmrun does not run these benchmarks.