Knowledge

HellaSwag Leaderboard

HellaSwag is a commonsense benchmark that asks a model to pick the most plausible continuation of an everyday situation. The wrong options are adversarially chosen to fool models, testing grounded commonsense understanding.

Source: epoch42 open models ranked+34 proprietaryData through Dec 2024

Open models ranked on HellaSwag

# shows rank among open models / rank overall (including proprietary).

#ModelScore
1 / 3Llama 3.1 405B · 405.9B
89.2%
2 / 4Falcon 180B · 180B
89.0%
3 / 5DeepSeek v3 · 684.5B
88.9%
4 / 6DeepSeek v2 · 235.7B
87.1%
5 / 8Mixtral 8x7B v0.1 · 46.7B
86.7%
6 / 10Llama 2 70B HF · 69.0B
85.3%
7 / 11Falcon 40B · 41.8B
85.3%
8 / 12Qwen2.5 72B · 72.7B
84.8%
9 / 14StableBeluga2 · 70B
84.1%
10 / 17Qwen2.5 Coder 32B · 32.8B
83.0%
11 / 18Falcon 11B · 11.1B
82.9%
12 / 23Gemma 7B · 8.5B
82.2%
13 / 26Mistral 7B v0.1 · 7B
81.0%
14 / 27Llama 2 13B HF · 13.0B
80.7%
15 / 28Qwen2.5 Coder 14B · 14.8B
80.2%
16 / 33Falcon 7B · 7.2B
78.1%
17 / 34Internlm 20B · 20B
78.1%
18 / 37Llama 2 7B HF · 6.7B
77.2%
19 / 38Phi 3 Small 8k Instruct · 7.4B
77.0%
20 / 39Qwen2.5 Coder 7B · 7.6B
76.8%
21 / 40Phi 3 Mini 4k Instruct · 3.8B
76.7%
22 / 42Yi 9B · 8.8B
76.4%
23 / 43Llama 7B · 6.7B
76.2%
24 / 45Bloom · 176.2B
74.4%
25 / 46Yi 6B · 6.1B
74.4%
26 / 47Xgen 7B 8k Base · 7B
74.2%
27 / 49INTELLECT 1 Instruct · 10.2B
71.4%
28 / 50Gemma 2B · 2.5B
71.4%
29 / 51Qwen2.5 Coder 3B · 3.1B
70.9%
30 / 52Baichuan2 13B Base · 13B
70.8%
31 / 54Internlm 7B · 7B
70.6%
32 / 55GPT Neox 20B · 20.7B
70.5%
33 / 59Baichuan2 7B Base · 7B
68.0%
34 / 61GPT J 6B · 6B
66.2%
35 / 62Qwen2.5 Coder 1.5B · 1.5B
61.8%
36 / 63Cerebras GPT 13B · 13B
59.4%
37 / 65Chatglm2 6B · 6B
57.0%
38 / 68Phi 2 · 2.8B
53.6%
39 / 69Qwen2.5 Coder 0.5B · 494M
48.4%
40 / 70Phi 1 5 · 1.4B
47.6%
41 / 75Stablelm Tuned Alpha 7B · 7B
40.7%
42 / 76Gpt2 Xl · 1.6B
40.0%

Score vs model size

Which models give the most quality for their size — the ones worth running locally.

1B10B100Bmodel size (log scale) →89.2%40.0%DeepSeek v3 · 685B · 88.9%DeepSeek v2 · 236B · 87.1%Llama 2 70B HF · 69B · 85.3%Qwen2.5 72B · 73B · 84.8%StableBeluga2 · 70B · 84.1%Llama 2 13B HF · 13B · 80.7%Qwen2.5 Coder 14B · 15B · 80.2%Internlm 20B · 20B · 78.1%Falcon 7B · 7B · 78.1%Phi 3 Small 8k Instruct · 7B · 77.0%Qwen2.5 Coder 7B · 8B · 76.8%Yi 9B · 9B · 76.4%Llama 7B · 7B · 76.2%Yi 6B · 6B · 74.4%Bloom · 176B · 74.4%Xgen 7B 8k Base · 7B · 74.2%INTELLECT 1 Instruct · 10B · 71.4%Qwen2.5 Coder 3B · 3B · 70.9%Baichuan2 13B Base · 13B · 70.8%Internlm 7B · 7B · 70.6%GPT Neox 20B · 21B · 70.5%Baichuan2 7B Base · 7B · 68.0%GPT J 6B · 6B · 66.2%Cerebras GPT 13B · 13B · 59.4%Chatglm2 6B · 6B · 57.0%Phi 2 · 3B · 53.6%Phi 1 5 · 1B · 47.6%Stablelm Tuned Alpha 7B · 7B · 40.7%Gpt2 Xl · 2B · 40.0%Qwen2.5 Coder 0.5B · 494M · 48.4%Qwen2.5 Coder 0.5BQwen2.5 Coder 1.5B · 2B · 61.8%Qwen2.5 Coder 1.5BGemma 2B · 3B · 71.4%Gemma 2BPhi 3 Mini 4k Instruct · 4B · 76.7%Llama 2 7B HF · 7B · 77.2%Llama 2 7B HFMistral 7B v0.1 · 7B · 81.0%Gemma 7B · 9B · 82.2%Gemma 7BFalcon 11B · 11B · 82.9%Falcon 11BQwen2.5 Coder 32B · 33B · 83.0%Falcon 40B · 42B · 85.3%Falcon 40BMixtral 8x7B v0.1 · 47B · 86.7%Mixtral 8x7B v0.1Falcon 180B · 180B · 89.0%Falcon 180BLlama 3.1 405B · 406B · 89.2%Llama 3.1 405B
Each dot is a model. Up = higher score, left = smaller (easier to run locally). The dashed line marks the efficiency frontier — the best score you can get at each size or smaller.
  • Qwen2.5 Coder 0.5B, 494M, score 48.4% — on the efficiency frontier (best score at its size or smaller).
  • Qwen2.5 Coder 1.5B, 2B, score 61.8% — on the efficiency frontier (best score at its size or smaller).
  • Gemma 2B, 3B, score 71.4% — on the efficiency frontier (best score at its size or smaller).
  • Phi 3 Mini 4k Instruct, 4B, score 76.7% — on the efficiency frontier (best score at its size or smaller).
  • Llama 2 7B HF, 7B, score 77.2% — on the efficiency frontier (best score at its size or smaller).
  • Mistral 7B v0.1, 7B, score 81.0% — on the efficiency frontier (best score at its size or smaller).
  • Gemma 7B, 9B, score 82.2% — on the efficiency frontier (best score at its size or smaller).
  • Falcon 11B, 11B, score 82.9% — on the efficiency frontier (best score at its size or smaller).
  • Qwen2.5 Coder 32B, 33B, score 83.0% — on the efficiency frontier (best score at its size or smaller).
  • Falcon 40B, 42B, score 85.3% — on the efficiency frontier (best score at its size or smaller).
  • Mixtral 8x7B v0.1, 47B, score 86.7% — on the efficiency frontier (best score at its size or smaller).
  • Falcon 180B, 180B, score 89.0% — on the efficiency frontier (best score at its size or smaller).
  • Llama 3.1 405B, 406B, score 89.2% — on the efficiency frontier (best score at its size or smaller).

HellaSwag: frequently asked questions

What is the best open LLM on HellaSwag?
Llama 3.1 405B is the top open model on HellaSwag, scoring 89.2%. Among all models tested — including proprietary ones — it ranks #3. The top model overall is GPT 4 (Mar 14) (OpenAI) at 95.3%.
What's the best HellaSwag model you can run on a 24 GB GPU?
Falcon 40B is the highest-scoring open model that fits in 24 GB at 4-bit quantization (about 23 GB), scoring 85.3% on HellaSwag.
What's the best HellaSwag model you can run on a 12 GB GPU?
Falcon 11B is the highest-scoring open model that fits in 12 GB at 4-bit quantization (about 6 GB), scoring 82.9% on HellaSwag.
Can open models match proprietary models on HellaSwag?
Not quite on HellaSwag: the strongest proprietary model (GPT 4 (Mar 14)) scores 95.3%, ahead of the best open model (Llama 3.1 405B) at 89.2% — but you can run the open one yourself.

Scores aggregated from epoch. llmrun does not run this benchmark — see the source for methodology, or the about benchmarks for what it measures.