Reasoning

WinoGrande Leaderboard

WinoGrande is a large-scale commonsense-reasoning test built around pronoun resolution: the model must pick which of two nouns a pronoun refers to in a sentence, using real-world knowledge to disambiguate. It was designed by AI2 to resist the shortcuts that let earlier models game its predecessor, Winograd.

Source: epoch46 open models ranked+34 proprietaryData through Dec 2024

Open models ranked on WinoGrande

# shows rank among open models / rank overall (including proprietary).

#ModelScore
1 / 1Llama 3.1 405B · 405.9B
89.2%
2 / 5Falcon 180B · 180B
87.1%
3 / 6DeepSeek v2 · 235.7B
86.3%
4 / 7DeepSeek v3 · 684.5B
85.2%
5 / 10Meta Llama 3 70B · 70.6B
83.5%
6 / 12Qwen2.5 72B · 72.7B
82.3%
7 / 13Llama 3.1 405B Instruct · 405.9B
82.2%
8 / 17Phi 3 Small 8k Instruct · 7.4B
81.5%
9 / 18Qwen2.5 Coder 32B · 32.8B
80.8%
10 / 19Llama 2 70B HF · 69.0B
80.2%
11 / 22Gemma 7B · 8.5B
79.0%
12 / 24Falcon 11B · 11.1B
78.3%
13 / 28Mixtral 8x7B v0.1 · 46.7B
77.2%
14 / 31Falcon 40B · 41.8B
76.9%
15 / 32Qwen2.5 Coder 14B · 14.8B
76.8%
16 / 35Meta Llama 3 8B · 8.0B
75.7%
17 / 36Mistral 7B v0.1 · 7B
75.3%
18 / 40Phi 1 5 · 1.4B
73.4%
19 / 42Yi 9B · 8.8B
73.0%
20 / 43DeepSeek Coder v2 Lite Base · 15.7B
72.9%
21 / 44Qwen2.5 Coder 7B · 7.6B
72.9%
22 / 45Llama 2 13B HF · 13.0B
72.8%
23 / 46Yi 6B · 6.1B
71.3%
24 / 48Phi 3 Mini 4k Instruct · 3.8B
70.8%
25 / 51Llama 7B · 6.7B
70.1%
26 / 52Llama 2 7B HF · 6.7B
69.2%
27 / 55Qwen2.5 Coder 3B · 3.1B
67.4%
28 / 56Falcon 7B · 7.2B
67.2%
29 / 58GPT Neox 20B · 20.7B
66.1%
30 / 59INTELLECT 1 Instruct · 10.2B
65.8%
31 / 60Gemma 2B · 2.5B
65.4%
32 / 61Meta Llama 3 8B Instruct · 8.0B
65.0%
33 / 62Xgen 7B 8k Base · 7B
64.9%
34 / 64GPT J 6B · 6B
64.5%
35 / 65Starcoder2 15B · 16.0B
64.3%
36 / 70Cerebras GPT 13B · 13B
60.8%
37 / 71Qwen2.5 Coder 1.5B · 1.5B
60.7%
38 / 72CodeQwen1.5 7B · 7.3B
59.8%
39 / 73Gpt2 Xl · 1.6B
58.3%
40 / 74Deepseek Coder 6.7B Base · 6.7B
57.6%
41 / 75Starcoder2 3B · 3.0B
57.1%
42 / 76Starcoder2 7B · 7.2B
57.1%
43 / 77Qwen2.5 Coder 0.5B · 494M
54.8%
44 / 78Phi 2 · 2.8B
54.7%
45 / 79Deepseek Coder 1.3B Base · 1.3B
53.3%
46 / 80Stablelm Tuned Alpha 7B · 7B
51.5%

Score vs model size

Which models give the most quality for their size — the ones worth running locally.

1B10B100Bmodel size (log scale) →89.2%51.5%DeepSeek v2 · 236B · 86.3%DeepSeek v3 · 685B · 85.2%Qwen2.5 72B · 73B · 82.3%Llama 3.1 405B Instruct · 406B · 82.2%Qwen2.5 Coder 32B · 33B · 80.8%Llama 2 70B HF · 69B · 80.2%Gemma 7B · 9B · 79.0%Falcon 11B · 11B · 78.3%Mixtral 8x7B v0.1 · 47B · 77.2%Falcon 40B · 42B · 76.9%Qwen2.5 Coder 14B · 15B · 76.8%Meta Llama 3 8B · 8B · 75.7%Yi 9B · 9B · 73.0%DeepSeek Coder v2 Lite Base · 16B · 72.9%Qwen2.5 Coder 7B · 8B · 72.9%Llama 2 13B HF · 13B · 72.8%Yi 6B · 6B · 71.3%Phi 3 Mini 4k Instruct · 4B · 70.8%Llama 7B · 7B · 70.1%Llama 2 7B HF · 7B · 69.2%Qwen2.5 Coder 3B · 3B · 67.4%Falcon 7B · 7B · 67.2%GPT Neox 20B · 21B · 66.1%INTELLECT 1 Instruct · 10B · 65.8%Gemma 2B · 3B · 65.4%Meta Llama 3 8B Instruct · 8B · 65.0%Xgen 7B 8k Base · 7B · 64.9%GPT J 6B · 6B · 64.5%Starcoder2 15B · 16B · 64.3%Cerebras GPT 13B · 13B · 60.8%Qwen2.5 Coder 1.5B · 2B · 60.7%CodeQwen1.5 7B · 7B · 59.8%Gpt2 Xl · 2B · 58.3%Deepseek Coder 6.7B Base · 7B · 57.6%Starcoder2 3B · 3B · 57.1%Starcoder2 7B · 7B · 57.1%Phi 2 · 3B · 54.7%Deepseek Coder 1.3B Base · 1B · 53.3%Stablelm Tuned Alpha 7B · 7B · 51.5%Qwen2.5 Coder 0.5B · 494M · 54.8%Qwen2.5 Coder 0.5BPhi 1 5 · 1B · 73.4%Phi 1 5Mistral 7B v0.1 · 7B · 75.3%Mistral 7B v0.1Phi 3 Small 8k Instruct · 7B · 81.5%Phi 3 Small 8k Instru…Meta Llama 3 70B · 71B · 83.5%Meta Llama 3 70BFalcon 180B · 180B · 87.1%Falcon 180BLlama 3.1 405B · 406B · 89.2%Llama 3.1 405B
Each dot is a model. Up = higher score, left = smaller (easier to run locally). The dashed line marks the efficiency frontier — the best score you can get at each size or smaller.
  • Qwen2.5 Coder 0.5B, 494M, score 54.8% — on the efficiency frontier (best score at its size or smaller).
  • Phi 1 5, 1B, score 73.4% — on the efficiency frontier (best score at its size or smaller).
  • Mistral 7B v0.1, 7B, score 75.3% — on the efficiency frontier (best score at its size or smaller).
  • Phi 3 Small 8k Instruct, 7B, score 81.5% — on the efficiency frontier (best score at its size or smaller).
  • Meta Llama 3 70B, 71B, score 83.5% — on the efficiency frontier (best score at its size or smaller).
  • Falcon 180B, 180B, score 87.1% — on the efficiency frontier (best score at its size or smaller).
  • Llama 3.1 405B, 406B, score 89.2% — on the efficiency frontier (best score at its size or smaller).

WinoGrande: frequently asked questions

What is the best open LLM on WinoGrande?
Llama 3.1 405B is the top open model on WinoGrande, scoring 89.2%. Among all models tested — including proprietary ones — it ranks #1. That puts it ahead of every proprietary model we track, including Claude 3 Opus (Feb 29, 2024) (Anthropic) at 88.5%.
What's the best WinoGrande model you can run on a 24 GB GPU?
Phi 3 Small 8k Instruct is the highest-scoring open model that fits in 24 GB at 4-bit quantization (about 4 GB), scoring 81.5% on WinoGrande.
What's the best WinoGrande model you can run on a 12 GB GPU?
Phi 3 Small 8k Instruct is the highest-scoring open model that fits in 12 GB at 4-bit quantization (about 4 GB), scoring 81.5% on WinoGrande.
Can open models match proprietary models on WinoGrande?
Yes — the best open model (Llama 3.1 405B, 89.2%) matches or beats every proprietary model we track on WinoGrande.

Scores aggregated from epoch. llmrun does not run this benchmark — see the source for methodology, or the about benchmarks for what it measures.