GPT OSS 20B — Benchmarks
Benchmark scores for GPT OSS 20B aggregated from public leaderboards, with how it ranks among open models. See hardware requirements for what you need to run it.
Overall rank: #67 of 123 open modelscomposite 44.9/100 across 6 benchmarks in 4 categories · methodology
Coding
| Benchmark | Score | Open rank | All models |
|---|---|---|---|
| Terminal-Bench | 3.4% | #16 / 16 | #57 / 57 |
| SciCode | 34.4% | #36 / 50 | #132 / 155 |
Knowledge
| Benchmark | Score | Open rank | All models |
|---|---|---|---|
| MMLU-Pro | 73.6% | #31 / 119 | #90 / 259 |
Math
| Benchmark | Score | Open rank | All models |
|---|---|---|---|
| AIME 2024/2025 | 65.3% | #30 / 74 | #139 / 270 |
Reasoning
| Benchmark | Score | Open rank | All models |
|---|---|---|---|
| GPQA Diamond | 60.8% | #38 / 83 | #189 / 291 |
| CritPt | 1.4% | #24 / 50 | #103 / 167 |
Scores aggregated from public benchmark sources (each linked from the benchmark pages). llmrun does not run these benchmarks.