Llama 2 70B HF — Benchmarks
Benchmark scores for Llama 2 70B HF aggregated from public leaderboards, with how it ranks among open models. See hardware requirements for what you need to run it.
Overall rank: #14 of 123 open modelscomposite 67.5/100 across 11 benchmarks in 3 categories · methodology
Knowledge
| Benchmark | Score | Open rank | All models |
|---|---|---|---|
| MMLU | 69.9% | #26 / 77 | #62 / 136 |
| HellaSwag | 85.3% | #6 / 42 | #10 / 76 |
| ARC Challenge | 78.3% | #14 / 51 | #20 / 77 |
| BoolQ | 88.6% | #3 / 33 | #9 / 77 |
| OpenBookQA | 60.2% | #9 / 21 | #15 / 42 |
| TriviaQA | 87.6% | #1 / 18 | #1 / 39 |
| MMLU-Pro | 37.5% | #87 / 119 | #208 / 259 |
Math
| Benchmark | Score | Open rank | All models |
|---|---|---|---|
| GSM8K | 69.6% | #17 / 60 | #26 / 93 |
Reasoning
| Benchmark | Score | Open rank | All models |
|---|---|---|---|
| BIG-Bench Hard | 64.9% | #9 / 37 | #13 / 50 |
| WinoGrande | 80.2% | #10 / 46 | #19 / 80 |
| PIQA | 82.8% | #10 / 35 | #17 / 60 |
Scores aggregated from public benchmark sources (each linked from the benchmark pages). llmrun does not run these benchmarks.