Llama 2 13B Chat HF — Benchmarks
Benchmark scores for Llama 2 13B Chat HF aggregated from public leaderboards, with how it ranks among open models. See hardware requirements for what you need to run it.
No overall rank for Llama 2 13B Chat HF: an overall score is only published when a model has been measured on enough of the benchmarks current models are still submitted to. Its individual scores below stand on their own — see the methodology for how the overall score is built.
Instruction Following
| Benchmark | Score | Open rank | All models |
|---|---|---|---|
| LMArena Text | 1141.2 | #166 / 182 | #359 / 391 |
Knowledge
| Benchmark | Score | Open rank | All models |
|---|---|---|---|
| MMLU | 50.9% | #62 / 85 | #108 / 136 |
Math
| Benchmark | Score | Open rank | All models |
|---|---|---|---|
| GSM8K | 36.9% | #38 / 62 | #54 / 93 |
Reasoning
| Benchmark | Score | Open rank | All models |
|---|---|---|---|
| Chess Puzzles | 0.0% | #53 / 65 | #193 / 205 |
| DTBench | 42.2% | #66 / 67 | #209 / 210 |
| BIG-Bench Hard | 58.2% | #13 / 38 | #19 / 50 |
Scores aggregated from public benchmark sources (each linked from the benchmark pages). llmrun does not run these benchmarks.