Llama 2 13B Chat HF — Benchmarks

Benchmark scores for Llama 2 13B Chat HF aggregated from public leaderboards, with how it ranks among open models. See hardware requirements for what you need to run it.

No overall rank for Llama 2 13B Chat HF: an overall score is only published when a model has been measured on enough of the benchmarks current models are still submitted to. Its individual scores below stand on their own — see the methodology for how the overall score is built.

Instruction Following

BenchmarkScoreOpen rankAll models
LMArena Text1141.2#166 / 182#359 / 391

Knowledge

BenchmarkScoreOpen rankAll models
MMLU50.9%#62 / 85#108 / 136

Math

BenchmarkScoreOpen rankAll models
GSM8K36.9%#38 / 62#54 / 93

Reasoning

BenchmarkScoreOpen rankAll models
Chess Puzzles0.0%#53 / 65#193 / 205
DTBench42.2%#66 / 67#209 / 210
BIG-Bench Hard58.2%#13 / 38#19 / 50

Scores aggregated from public benchmark sources (each linked from the benchmark pages). llmrun does not run these benchmarks.