Llama 7B — Benchmarks
Benchmark scores for Llama 7B aggregated from public leaderboards, with how it ranks among open models. See hardware requirements for what you need to run it.
Overall rank: #86 of 123 open modelscomposite 37.1/100 across 10 benchmarks in 3 categories · methodology
Knowledge
| Benchmark | Score | Open rank | All models |
|---|---|---|---|
| HellaSwag | 76.2% | #23 / 42 | #43 / 76 |
| ARC Challenge | 47.6% | #32 / 51 | #50 / 77 |
| BoolQ | 76.5% | #18 / 33 | #46 / 77 |
| OpenBookQA | 57.2% | #11 / 21 | #20 / 42 |
| TriviaQA | 71.0% | #12 / 18 | #30 / 39 |
| MMLU | 35.6% | #72 / 77 | #126 / 136 |
Math
| Benchmark | Score | Open rank | All models |
|---|---|---|---|
| GSM8K | 11.0% | #56 / 60 | #79 / 93 |
Reasoning
| Benchmark | Score | Open rank | All models |
|---|---|---|---|
| BIG-Bench Hard | 33.5% | #34 / 37 | #47 / 50 |
| WinoGrande | 70.1% | #25 / 46 | #51 / 80 |
| PIQA | 79.8% | #21 / 35 | #39 / 60 |
Scores aggregated from public benchmark sources (each linked from the benchmark pages). llmrun does not run these benchmarks.