Reasoning

WinoGrande Leaderboard

WinoGrande is a large-scale commonsense-reasoning test built around pronoun resolution: the model must pick which of two nouns a pronoun refers to in a sentence, using real-world knowledge to disambiguate. It was designed by AI2 to resist the shortcuts that let earlier models game its predecessor, Winograd.

Source: epoch46 open models ranked+34 proprietaryData through Dec 2024

All models ranked on WinoGrande

Proprietary / closed models are shown dimmed — you can't run them locally, but they show where the open field stands.

#ModelScore
1Llama 3.1 405B · 405.9B
89.2%
2Claude 3 Opus (Feb 29, 2024) · proprietary
88.5%
3GPT 4 (Mar 14) · proprietary
87.5%
4GPT 4 32K (Mar 14) · proprietary
87.5%
5Falcon 180B · 180B
87.1%
6DeepSeek v2 · 235.7B
86.3%
7DeepSeek v3 · 684.5B
85.2%
8PaLM 540B · proprietary
85.1%
9DeepSeek Coder V2 Base · proprietary
83.7%
10Meta Llama 3 70B · 70.6B
83.5%
11PaLM 2 L · proprietary
83.0%
12Qwen2.5 72B · 72.7B
82.3%
13Llama 3.1 405B Instruct · 405.9B
82.2%
14GPT 3.5 Turbo (Jun 13) · proprietary
81.6%
15Text Davinci 002 · proprietary
81.6%
16Phi 3 Medium 128K Instruct · proprietary
81.5%
17Phi 3 Small 8k Instruct · 7.4B
81.5%
18Qwen2.5 Coder 32B · 32.8B
80.8%
19Llama 2 70B HF · 69.0B
80.2%
20GLaM (MoE) · proprietary
79.2%
21PaLM 2 M · proprietary
79.2%
22Gemma 7B · 8.5B
79.0%
23Megatron Turing NLG 530B · proprietary
78.9%
24Falcon 11B · 11.1B
78.3%
25Nemotron 4 15B · proprietary
78.0%
26PaLM 2 S · proprietary
77.9%
27Text Davinci 001 · proprietary
77.7%
28Mixtral 8x7B v0.1 · 46.7B
77.2%
29Llama 65B · proprietary
77.0%
30PaLM 62B · proprietary
77.0%
31Falcon 40B · 41.8B
76.9%
32Qwen2.5 Coder 14B · 14.8B
76.8%
33Llama 2 34B · proprietary
76.7%
34Llama 33B · proprietary
76.0%
35Meta Llama 3 8B · 8.0B
75.7%
36Mistral 7B v0.1 · 7B
75.3%
37Claude 3 Sonnet (Feb 29, 2024) · proprietary
75.1%
38Chinchilla (70B) · proprietary
74.9%
39Claude 3 Haiku (Mar 07, 2024) · proprietary
74.2%
40Phi 1 5 · 1.4B
73.4%
41Llama 13B · proprietary
73.0%
42Yi 9B · 8.8B
73.0%
43DeepSeek Coder v2 Lite Base · 15.7B
72.9%
44Qwen2.5 Coder 7B · 7.6B
72.9%
45Llama 2 13B HF · 13.0B
72.8%
46Yi 6B · 6.1B
71.3%
47Mpt 30B · proprietary
71.0%
48Phi 3 Mini 4k Instruct · 3.8B
70.8%
49Vicuna 13B v1.1 · proprietary
70.8%
50Gopher (280B) · proprietary
70.2%
51Llama 7B · 6.7B
70.1%
52Llama 2 7B HF · 6.7B
69.2%
53GPT 3.5 Turbo (Nov 06) · proprietary
68.8%
54Mpt 7B · proprietary
68.6%
55Qwen2.5 Coder 3B · 3.1B
67.4%
56Falcon 7B · 7.2B
67.2%
57Open Llama 7B · proprietary
67.0%
58GPT Neox 20B · 20.7B
66.1%
59INTELLECT 1 Instruct · 10.2B
65.8%
60Gemma 2B · 2.5B
65.4%
61Meta Llama 3 8B Instruct · 8.0B
65.0%
62Xgen 7B 8k Base · 7B
64.9%
63Opt 13B · proprietary
64.7%
64GPT J 6B · 6B
64.5%
65Starcoder2 15B · 16.0B
64.3%
66RedPajama INCITE 7B Base · proprietary
63.8%
67DeepSeek Coder 33B Base · proprietary
62.0%
68Dolly v2 12B · proprietary
61.8%
69Opt 1.3b · proprietary
61.0%
70Cerebras GPT 13B · 13B
60.8%
71Qwen2.5 Coder 1.5B · 1.5B
60.7%
72CodeQwen1.5 7B · 7.3B
59.8%
73Gpt2 Xl · 1.6B
58.3%
74Deepseek Coder 6.7B Base · 6.7B
57.6%
75Starcoder2 3B · 3.0B
57.1%
76Starcoder2 7B · 7.2B
57.1%
77Qwen2.5 Coder 0.5B · 494M
54.8%
78Phi 2 · 2.8B
54.7%
79Deepseek Coder 1.3B Base · 1.3B
53.3%
80Stablelm Tuned Alpha 7B · 7B
51.5%

Score vs model size

Which models give the most quality for their size — the ones worth running locally.

1B10B100Bmodel size (log scale) →89.2%51.5%DeepSeek v2 · 236B · 86.3%DeepSeek v3 · 685B · 85.2%Qwen2.5 72B · 73B · 82.3%Llama 3.1 405B Instruct · 406B · 82.2%Qwen2.5 Coder 32B · 33B · 80.8%Llama 2 70B HF · 69B · 80.2%Gemma 7B · 9B · 79.0%Falcon 11B · 11B · 78.3%Mixtral 8x7B v0.1 · 47B · 77.2%Falcon 40B · 42B · 76.9%Qwen2.5 Coder 14B · 15B · 76.8%Meta Llama 3 8B · 8B · 75.7%Yi 9B · 9B · 73.0%DeepSeek Coder v2 Lite Base · 16B · 72.9%Qwen2.5 Coder 7B · 8B · 72.9%Llama 2 13B HF · 13B · 72.8%Yi 6B · 6B · 71.3%Phi 3 Mini 4k Instruct · 4B · 70.8%Llama 7B · 7B · 70.1%Llama 2 7B HF · 7B · 69.2%Qwen2.5 Coder 3B · 3B · 67.4%Falcon 7B · 7B · 67.2%GPT Neox 20B · 21B · 66.1%INTELLECT 1 Instruct · 10B · 65.8%Gemma 2B · 3B · 65.4%Meta Llama 3 8B Instruct · 8B · 65.0%Xgen 7B 8k Base · 7B · 64.9%GPT J 6B · 6B · 64.5%Starcoder2 15B · 16B · 64.3%Cerebras GPT 13B · 13B · 60.8%Qwen2.5 Coder 1.5B · 2B · 60.7%CodeQwen1.5 7B · 7B · 59.8%Gpt2 Xl · 2B · 58.3%Deepseek Coder 6.7B Base · 7B · 57.6%Starcoder2 3B · 3B · 57.1%Starcoder2 7B · 7B · 57.1%Phi 2 · 3B · 54.7%Deepseek Coder 1.3B Base · 1B · 53.3%Stablelm Tuned Alpha 7B · 7B · 51.5%Qwen2.5 Coder 0.5B · 494M · 54.8%Qwen2.5 Coder 0.5BPhi 1 5 · 1B · 73.4%Phi 1 5Mistral 7B v0.1 · 7B · 75.3%Mistral 7B v0.1Phi 3 Small 8k Instruct · 7B · 81.5%Phi 3 Small 8k Instru…Meta Llama 3 70B · 71B · 83.5%Meta Llama 3 70BFalcon 180B · 180B · 87.1%Falcon 180BLlama 3.1 405B · 406B · 89.2%Llama 3.1 405B
Each dot is a model. Up = higher score, left = smaller (easier to run locally). The dashed line marks the efficiency frontier — the best score you can get at each size or smaller.
  • Qwen2.5 Coder 0.5B, 494M, score 54.8% — on the efficiency frontier (best score at its size or smaller).
  • Phi 1 5, 1B, score 73.4% — on the efficiency frontier (best score at its size or smaller).
  • Mistral 7B v0.1, 7B, score 75.3% — on the efficiency frontier (best score at its size or smaller).
  • Phi 3 Small 8k Instruct, 7B, score 81.5% — on the efficiency frontier (best score at its size or smaller).
  • Meta Llama 3 70B, 71B, score 83.5% — on the efficiency frontier (best score at its size or smaller).
  • Falcon 180B, 180B, score 87.1% — on the efficiency frontier (best score at its size or smaller).
  • Llama 3.1 405B, 406B, score 89.2% — on the efficiency frontier (best score at its size or smaller).

WinoGrande: frequently asked questions

What is the best open LLM on WinoGrande?
Llama 3.1 405B is the top open model on WinoGrande, scoring 89.2%. Among all models tested — including proprietary ones — it ranks #1. That puts it ahead of every proprietary model we track, including Claude 3 Opus (Feb 29, 2024) (Anthropic) at 88.5%.
What's the best WinoGrande model you can run on a 24 GB GPU?
Phi 3 Small 8k Instruct is the highest-scoring open model that fits in 24 GB at 4-bit quantization (about 4 GB), scoring 81.5% on WinoGrande.
What's the best WinoGrande model you can run on a 12 GB GPU?
Phi 3 Small 8k Instruct is the highest-scoring open model that fits in 12 GB at 4-bit quantization (about 4 GB), scoring 81.5% on WinoGrande.
Can open models match proprietary models on WinoGrande?
Yes — the best open model (Llama 3.1 405B, 89.2%) matches or beats every proprietary model we track on WinoGrande.

Scores aggregated from epoch. llmrun does not run this benchmark — see the source for methodology, or the about benchmarks for what it measures.