Llama 2 70B HF — Benchmarks

Benchmark scores for Llama 2 70B HF aggregated from public leaderboards, with how it ranks among open models. See hardware requirements for what you need to run it.

Overall rank: #14 of 123 open modelscomposite 67.5/100 across 11 benchmarks in 3 categories · methodology

Knowledge

BenchmarkScoreOpen rankAll models
MMLU69.9%#26 / 77#62 / 136
HellaSwag85.3%#6 / 42#10 / 76
ARC Challenge78.3%#14 / 51#20 / 77
BoolQ88.6%#3 / 33#9 / 77
OpenBookQA60.2%#9 / 21#15 / 42
TriviaQA87.6%#1 / 18#1 / 39
MMLU-Pro37.5%#87 / 119#208 / 259

Math

BenchmarkScoreOpen rankAll models
GSM8K69.6%#17 / 60#26 / 93

Reasoning

BenchmarkScoreOpen rankAll models
BIG-Bench Hard64.9%#9 / 37#13 / 50
WinoGrande80.2%#10 / 46#19 / 80
PIQA82.8%#10 / 35#17 / 60

Scores aggregated from public benchmark sources (each linked from the benchmark pages). llmrun does not run these benchmarks.

Llama 2 70B HF Benchmarks — Scores & Rankings | llmrun