Gpt2 Xl — Benchmarks
Benchmark scores for Gpt2 Xl aggregated from public leaderboards, with how it ranks among open models. See hardware requirements for what you need to run it.
Knowledge
| Benchmark | Score | Open rank | All models |
|---|---|---|---|
| HellaSwag | 40.0% | #42 / 42 | #76 / 76 |
| ARC Challenge | 25.0% | #51 / 51 | #76 / 77 |
| BoolQ | 61.8% | #31 / 33 | #69 / 77 |
| OpenBookQA | 22.4% | #21 / 21 | #42 / 42 |
Reasoning
| Benchmark | Score | Open rank | All models |
|---|---|---|---|
| WinoGrande | 58.3% | #39 / 46 | #73 / 80 |
| PIQA | 70.5% | #33 / 35 | #57 / 60 |
Scores aggregated from public benchmark sources (each linked from the benchmark pages). llmrun does not run these benchmarks.