Llama 2 70B Chat HF — Benchmarks

Benchmark scores for Llama 2 70B Chat HF aggregated from public leaderboards, with how it ranks among open models. See hardware requirements for what you need to run it.

Overall rank: #57 of 73 open modelscomposite 32.4/100 across 6 benchmarks in 3 categories · methodology

Knowledge

BenchmarkScoreOpen rankAll models
MMLU59.9%#43 / 76#90 / 136

Math

BenchmarkScoreOpen rankAll models
AIME 2024/20250.0%#34 / 34#155 / 155
MATH Level 53.3%#32 / 32#108 / 108
GSM8K58.7%#25 / 59#34 / 93

Reasoning

BenchmarkScoreOpen rankAll models
GPQA Diamond26.3%#41 / 46#175 / 182
BIG-Bench Hard58.5%#11 / 37#17 / 50

Scores aggregated from public benchmark sources (each linked from the benchmark pages). llmrun does not run these benchmarks.