Llama 2 13B HF — Benchmarks

Benchmark scores for Llama 2 13B HF aggregated from public leaderboards, with how it ranks among open models. See hardware requirements for what you need to run it.

No overall rank for Llama 2 13B HF: an overall score is only published when a model has been measured on enough of the benchmarks current models are still submitted to. Its individual scores below stand on their own — see the methodology for how the overall score is built.

Knowledge

BenchmarkScoreOpen rankAll models
MMLU-Pro25.3%#125 / 144#234 / 257
ARC Challenge60.3%#25 / 56#35 / 76
BoolQ82.4%#14 / 38#35 / 76
OpenBookQA57.0%#13 / 26#21 / 41
TriviaQA79.6%#7 / 19#17 / 39
MMLU55.6%#56 / 85#101 / 136
HellaSwag80.7%#15 / 47#27 / 75

Math

BenchmarkScoreOpen rankAll models
GSM8K34.3%#42 / 62#59 / 93

Reasoning

BenchmarkScoreOpen rankAll models
WinoGrande72.8%#23 / 51#45 / 79
BIG-Bench Hard47.0%#23 / 38#30 / 50
PIQA80.8%#18 / 40#32 / 59

Scores aggregated from public benchmark sources (each linked from the benchmark pages). llmrun does not run these benchmarks.