Llama 2 13B HF — Benchmarks
Benchmark scores for Llama 2 13B HF aggregated from public leaderboards, with how it ranks among open models. See hardware requirements for what you need to run it.
No overall rank for Llama 2 13B HF: an overall score is only published when a model has been measured on enough of the benchmarks current models are still submitted to. Its individual scores below stand on their own — see the methodology for how the overall score is built.
Knowledge
| Benchmark | Score | Open rank | All models |
|---|---|---|---|
| MMLU-Pro | 25.3% | #125 / 144 | #234 / 257 |
| ARC Challenge | 60.3% | #25 / 56 | #35 / 76 |
| BoolQ | 82.4% | #14 / 38 | #35 / 76 |
| OpenBookQA | 57.0% | #13 / 26 | #21 / 41 |
| TriviaQA | 79.6% | #7 / 19 | #17 / 39 |
| MMLU | 55.6% | #56 / 85 | #101 / 136 |
| HellaSwag | 80.7% | #15 / 47 | #27 / 75 |
Math
| Benchmark | Score | Open rank | All models |
|---|---|---|---|
| GSM8K | 34.3% | #42 / 62 | #59 / 93 |
Reasoning
| Benchmark | Score | Open rank | All models |
|---|---|---|---|
| WinoGrande | 72.8% | #23 / 51 | #45 / 79 |
| BIG-Bench Hard | 47.0% | #23 / 38 | #30 / 50 |
| PIQA | 80.8% | #18 / 40 | #32 / 59 |
Scores aggregated from public benchmark sources (each linked from the benchmark pages). llmrun does not run these benchmarks.