Knowledge

MMLU Leaderboard

MMLU (Massive Multitask Language Understanding) spans 57 subjects from history to law to medicine as multiple-choice questions. It is the long-standing default for broad knowledge and remains the most widely-reported general benchmark.

Source: epoch76 open models ranked+60 proprietaryData through Feb 2025

All models ranked on MMLU

Proprietary / closed models are shown dimmed — you can't run them locally, but they show where the open field stands.

#ModelScore
1GPT 4o (Nov 20, 2024) · proprietary
88.1%
2Claude 3.5 Sonnet (Oct 22, 2024) · proprietary
87.3%
3DeepSeek v3 · 684.5B
87.2%
4Gemini 1.5 Pro 002 · proprietary
86.9%
5Claude 3.5 Sonnet (Jun 20, 2024) · proprietary
86.5%
6GPT 4 (Mar 14) · proprietary
86.4%
7Llama 3.3 70B Instruct · 70.6B
86.3%
8Gemini 1.5 Pro 001 · proprietary
85.9%
9Qwen2.5 72B Instruct · 72.7B
85.3%
10Qwen2.5 72B · 72.7B
85.0%
11Phi 4 · 14.7B
84.8%
12Claude 3 Opus (Feb 29, 2024) · proprietary
84.6%
13Llama 3.1 405B Instruct · 405.9B
84.5%
14Llama 3.1 405B · 405.9B
84.4%
15GPT 4o (Aug 06, 2024) · proprietary
84.3%
16GPT 4o (May 13, 2024) · proprietary
84.2%
17Gemini 1.5 Pro 001 Feb24 · proprietary
82.7%
18GPT 4 (Jun 13) · proprietary
82.4%
19Qwen2 72B Instruct · 72.7B
82.4%
20Amazon.nova Pro v1:0 · proprietary
82.0%
21GPT 4o Mini (Jul 18, 2024) · proprietary
81.8%
22GPT 4 Turbo (Apr 09, 2024) · proprietary
81.3%
23Llama 3.2 90B Vision Instruct · 88.6B
80.3%
24Llama 3.1 70B Instruct · 70.6B
80.1%
25Mistral Large 2407 · proprietary
80.0%
26Qwen2.5 14B Instruct · 14.8B
79.9%
27Gemini 2.0 Flash Exp · proprietary
79.7%
28GPT 4 Turbo · proprietary
79.6%
29Meta Llama 3 70B Instruct · 70.6B
79.3%
30Yi Large · proprietary
79.3%
31Qwen2.5 Coder 32B · 32.8B
79.1%
32Claude 2.0 · proprietary
78.5%
33DeepSeek v2 · 235.7B
78.4%
34Phi 3 Medium 128K Instruct · proprietary
78.0%
35Gemini 1.5 Flash 001 · proprietary
77.9%
36Gemini 1.5 Flash (May 14) · proprietary
77.8%
37Mixtral 8x22B v0.1 · 140.6B
77.8%
38Amazon.nova Lite v1:0 · proprietary
77.0%
39Claude 1.3 · proprietary
77.0%
40Yi 34B · 34.4B
76.3%
41Claude 3 Sonnet (Feb 29, 2024) · proprietary
75.9%
42Gemma 2 27B IT · 27.2B
75.7%
43Phi 3 Small 8k Instruct · 7.4B
75.7%
44Qwen2.5 Coder 14B · 14.8B
75.2%
45Qwen1.5 32B · 32.5B
74.4%
46Claude 3.5 Haiku (Oct 22, 2024) · proprietary
74.3%
47Gemini 1.5 Flash 002 · proprietary
73.9%
48Claude 3 Haiku (Mar 07, 2024) · proprietary
73.8%
49Claude 2.1 · proprietary
73.5%
50Yi 34B Chat · 34.4B
73.5%
51Claude Instant 1.1 · proprietary
73.4%
52Claude Instant 1.2 · proprietary
73.2%
53Qwen2.5 7B Instruct · 7.6B
72.9%
54Inflection 1 · proprietary
72.7%
55Gemma 2 9B IT · 9.2B
72.1%
56GPT 3.5 Turbo (Nov 06) · proprietary
71.4%
57Amazon.nova Micro v1:0 · proprietary
70.8%
58Falcon 180B · 180B
70.6%
59Mixtral 8x7B v0.1 · 46.7B
70.6%
60Gemini 1.0 Pro 001 · proprietary
70.0%
61Text Davinci 002 · proprietary
70.0%
62Llama 2 70B HF · 69.0B
69.9%
63C4ai Command R Plus (Aug 2024) · proprietary
69.4%
64PaLM 540B · proprietary
69.3%
65GPT 3.5 Turbo (Jun 13) · proprietary
68.9%
66Meta Llama 3 8B Instruct · 8.0B
68.8%
67Mistral Large 2402 · proprietary
68.8%
68Phi 3 Mini 4k Instruct · 3.8B
68.8%
69Mistral Small 2402 · proprietary
68.7%
70Qwen1.5 14B · 14.2B
68.6%
71StableBeluga2 · 70B
68.6%
72Yi 9B · 8.8B
68.4%
73Qwen2.5 Coder 7B · 7.6B
68.0%
74Chinchilla (70B) · proprietary
67.5%
75GPT 3.5 Turbo (Jan 25) · proprietary
67.3%
76Qwen 14B · 14.2B
66.3%
77Gemma 7B · 8.5B
66.1%
78C4ai Command R (Aug 2024) · proprietary
65.2%
79Qwen 14B Chat · 14.2B
65.0%
80Starcoder2 15B · 16.0B
64.1%
81Yi 6B · 6.1B
64.0%
82Llama 65B · proprietary
63.4%
83Llama 2 34B · proprietary
62.6%
84Qwen1.5 7B · 7.7B
62.6%
85Mistral 7B Instruct v0.2 · 7.2B
62.5%
86Mistral 7B v0.1 · 7B
62.5%
87Yi 6B Chat · 6.1B
61.0%
88DeepSeek Coder v2 Lite Base · 15.7B
60.5%
89Gopher (280B) · proprietary
60.0%
90Llama 2 70B Chat HF · 69.0B
59.9%
91Mistral 7B Instruct v0.3 · 7.2B
59.9%
92Baichuan2 13B Base · 13B
59.2%
93Llama 33B · proprietary
58.7%
94Nemotron 4 15B · proprietary
58.7%
95Falcon 11B · 11.1B
58.4%
96Phi 2 · 2.8B
58.4%
97Internlm Chat 20B · 20B
57.4%
98Falcon 40B · 41.8B
56.9%
99Llama 3.2 11B Vision Instruct · 10.7B
56.5%
100Llama 3.1 8B Instruct · 8.0B
56.1%
101Llama 2 13B HF · 13.0B
55.6%
102Baichuan2 13B Chat · 13B
55.1%
103Baichuan2 7B Base · 7B
54.2%
104PaLM 62B · proprietary
53.7%
105Qwen2.5 Coder 1.5B · 1.5B
53.6%
106Baichuan 13B Base · 13B
51.6%
107Internlm 7B · 7B
51.0%
108Llama 2 13B Chat HF · 13.0B
50.9%
109INTELLECT 1 Instruct · 10.2B
49.9%
110Chatglm2 6B · 6B
47.9%
111Mpt 30B · proprietary
47.9%
112Llama 13B · proprietary
47.7%
113Llama 2 7B HF · 6.7B
45.8%
114Qwen 7B · 7.7B
45.0%
115Text Davinci 001 · proprietary
43.9%
116Baichuan 7B · 7B
42.3%
117Gemma 2B · 2.5B
42.3%
118Qwen2.5 Coder 0.5B · 494M
42.0%
119CodeQwen1.5 7B · 7.3B
40.5%
120DeepSeek Coder 33B Base · proprietary
39.4%
121Starcoder2 7B · 7.2B
38.8%
122Phi 1 5 · 1.4B
37.6%
123Starcoder2 3B · 3.0B
36.6%
124DeepSeek Coder 6.7b Base · proprietary
36.4%
125Xgen 7B 8k Base · 7B
36.3%
126Llama 7B · 6.7B
35.6%
127Falcon 7B · 7.2B
35.0%
128Mpt 7B · proprietary
30.8%
129Open Llama 7B · proprietary
29.9%
130Qwen 1 8B · 1.8B
28.2%
131RedPajama INCITE 7B Base · proprietary
26.3%
132Cerebras GPT 13B · 13B
26.2%
133Dolly v2 12B · proprietary
26.2%
134Deepseek Coder 1.3B Base · 1.3B
25.8%
135GPT J 6B · 6B
25.7%
136Opt 13B · proprietary
25.1%

Score vs model size

Which models give the most quality for their size — the ones worth running locally.

1B10B100Bmodel size (log scale) →87.2%25.7%Qwen2.5 72B Instruct · 73B · 85.3%Qwen2.5 72B · 73B · 85.0%Llama 3.1 405B Instruct · 406B · 84.5%Llama 3.1 405B · 406B · 84.4%Qwen2 72B Instruct · 73B · 82.4%Llama 3.2 90B Vision Instruct · 89B · 80.3%Llama 3.1 70B Instruct · 71B · 80.1%Qwen2.5 14B Instruct · 15B · 79.9%Meta Llama 3 70B Instruct · 71B · 79.3%Qwen2.5 Coder 32B · 33B · 79.1%DeepSeek v2 · 236B · 78.4%Mixtral 8x22B v0.1 · 141B · 77.8%Yi 34B · 34B · 76.3%Gemma 2 27B IT · 27B · 75.7%Qwen2.5 Coder 14B · 15B · 75.2%Qwen1.5 32B · 33B · 74.4%Yi 34B Chat · 34B · 73.5%Qwen2.5 7B Instruct · 8B · 72.9%Gemma 2 9B IT · 9B · 72.1%Mixtral 8x7B v0.1 · 47B · 70.6%Falcon 180B · 180B · 70.6%Llama 2 70B HF · 69B · 69.9%Meta Llama 3 8B Instruct · 8B · 68.8%Qwen1.5 14B · 14B · 68.6%StableBeluga2 · 70B · 68.6%Yi 9B · 9B · 68.4%Qwen2.5 Coder 7B · 8B · 68.0%Qwen 14B · 14B · 66.3%Gemma 7B · 9B · 66.1%Qwen 14B Chat · 14B · 65.0%Starcoder2 15B · 16B · 64.1%Yi 6B · 6B · 64.0%Qwen1.5 7B · 8B · 62.6%Mistral 7B Instruct v0.2 · 7B · 62.5%Mistral 7B v0.1 · 7B · 62.5%Yi 6B Chat · 6B · 61.0%DeepSeek Coder v2 Lite Base · 16B · 60.5%Llama 2 70B Chat HF · 69B · 59.9%Mistral 7B Instruct v0.3 · 7B · 59.9%Baichuan2 13B Base · 13B · 59.2%Falcon 11B · 11B · 58.4%Internlm Chat 20B · 20B · 57.4%Falcon 40B · 42B · 56.9%Llama 3.2 11B Vision Instruct · 11B · 56.5%Llama 3.1 8B Instruct · 8B · 56.1%Llama 2 13B HF · 13B · 55.6%Baichuan2 13B Chat · 13B · 55.1%Baichuan2 7B Base · 7B · 54.2%Baichuan 13B Base · 13B · 51.6%Internlm 7B · 7B · 51.0%Llama 2 13B Chat HF · 13B · 50.9%INTELLECT 1 Instruct · 10B · 49.9%Chatglm2 6B · 6B · 47.9%Llama 2 7B HF · 7B · 45.8%Qwen 7B · 8B · 45.0%Baichuan 7B · 7B · 42.3%Gemma 2B · 3B · 42.3%CodeQwen1.5 7B · 7B · 40.5%Starcoder2 7B · 7B · 38.8%Phi 1 5 · 1B · 37.6%Starcoder2 3B · 3B · 36.6%Xgen 7B 8k Base · 7B · 36.3%Llama 7B · 7B · 35.6%Falcon 7B · 7B · 35.0%Qwen 1 8B · 2B · 28.2%Cerebras GPT 13B · 13B · 26.2%Deepseek Coder 1.3B Base · 1B · 25.8%GPT J 6B · 6B · 25.7%Qwen2.5 Coder 0.5B · 494M · 42.0%Qwen2.5 Coder 0.5BQwen2.5 Coder 1.5B · 2B · 53.6%Qwen2.5 Coder 1.5BPhi 2 · 3B · 58.4%Phi 2Phi 3 Mini 4k Instruct · 4B · 68.8%Phi 3 Mini 4k InstructPhi 3 Small 8k Instruct · 7B · 75.7%Phi 3 Small 8k Instru…Phi 4 · 15B · 84.8%Phi 4Llama 3.3 70B Instruct · 71B · 86.3%Llama 3.3 70B InstructDeepSeek v3 · 685B · 87.2%DeepSeek v3
Each dot is a model. Up = higher score, left = smaller (easier to run locally). The dashed line marks the efficiency frontier — the best score you can get at each size or smaller.
  • Qwen2.5 Coder 0.5B, 494M, score 42.0% — on the efficiency frontier (best score at its size or smaller).
  • Qwen2.5 Coder 1.5B, 2B, score 53.6% — on the efficiency frontier (best score at its size or smaller).
  • Phi 2, 3B, score 58.4% — on the efficiency frontier (best score at its size or smaller).
  • Phi 3 Mini 4k Instruct, 4B, score 68.8% — on the efficiency frontier (best score at its size or smaller).
  • Phi 3 Small 8k Instruct, 7B, score 75.7% — on the efficiency frontier (best score at its size or smaller).
  • Phi 4, 15B, score 84.8% — on the efficiency frontier (best score at its size or smaller).
  • Llama 3.3 70B Instruct, 71B, score 86.3% — on the efficiency frontier (best score at its size or smaller).
  • DeepSeek v3, 685B, score 87.2% — on the efficiency frontier (best score at its size or smaller).

MMLU: frequently asked questions

What is the best open LLM on MMLU?
DeepSeek v3 is the top open model on MMLU, scoring 87.2%. Among all models tested — including proprietary ones — it ranks #3. The top model overall is GPT 4o (Nov 20, 2024) (OpenAI) at 88.1%.
What's the best MMLU model you can run on a 24 GB GPU?
Phi 4 is the highest-scoring open model that fits in 24 GB at 4-bit quantization (about 8 GB), scoring 84.8% on MMLU.
What's the best MMLU model you can run on a 12 GB GPU?
Phi 4 is the highest-scoring open model that fits in 12 GB at 4-bit quantization (about 8 GB), scoring 84.8% on MMLU.
Can open models match proprietary models on MMLU?
Not quite on MMLU: the strongest proprietary model (GPT 4o (Nov 20, 2024)) scores 88.1%, ahead of the best open model (DeepSeek v3) at 87.2% — but you can run the open one yourself.

Scores aggregated from epoch. llmrun does not run this benchmark — see the source for methodology, or the about benchmarks for what it measures.