Llama 2 7B HF — Benchmarks
Benchmark scores for Llama 2 7B HF aggregated from public leaderboards, with how it ranks among open models. See hardware requirements for what you need to run it.
No overall rank for Llama 2 7B HF: an overall score is only published when a model has been measured on enough of the benchmarks current models are still submitted to. Its individual scores below stand on their own — see the methodology for how the overall score is built.
Knowledge
| Benchmark | Score | Open rank | All models |
|---|---|---|---|
| MMLU-Pro | 20.3% | #130 / 144 | #242 / 257 |
| ARC Challenge | 45.9% | #35 / 56 | #52 / 76 |
| BoolQ | 77.9% | #18 / 38 | #44 / 76 |
| OpenBookQA | 58.6% | #11 / 26 | #17 / 41 |
| TriviaQA | 73.7% | #10 / 19 | #25 / 39 |
| MMLU | 45.8% | #65 / 85 | #113 / 136 |
| HellaSwag | 77.2% | #19 / 47 | #37 / 75 |
Math
| Benchmark | Score | Open rank | All models |
|---|---|---|---|
| GSM8K | 16.7% | #56 / 62 | #76 / 93 |
Reasoning
| Benchmark | Score | Open rank | All models |
|---|---|---|---|
| WinoGrande | 69.2% | #27 / 51 | #52 / 79 |
| BIG-Bench Hard | 39.2% | #29 / 38 | #38 / 50 |
| PIQA | 78.8% | #23 / 40 | #40 / 59 |
Scores aggregated from public benchmark sources (each linked from the benchmark pages). llmrun does not run these benchmarks.