Knowledge

MMLU-Pro Leaderboard

MMLU-Pro is a harder, cleaned-up successor to MMLU with ten answer choices and more reasoning-heavy questions across 14 subjects, measuring broad knowledge and reasoning together.

Source: tigerlab144 open models ranked+113 proprietary

Open models ranked on MMLU-Pro

# shows rank among open models / rank overall (including proprietary).

#ModelScore
1 / 6MiniMax M2.1 · 228.7B
88.0%
2 / 7Qwen3.5 397B A17B · 403.4B
87.8%
3 / 17Kimi K2.5 · 1026.9B
87.1%
4 / 21Qwen3.5 122B A10B · 125.1B
86.7%
5 / 28Qwen3.5 27B · 27.8B
86.1%
6 / 30GLM 5 · 753.9B
86.0%
7 / 32Qwen3.5 35B A3B · 36.0B
85.3%
8 / 33DeepSeek V3.2 · 685.4B
85.0%
9 / 35DeepSeek V3.1 · 684.5B
84.8%
10 / 36GLM 4.5 · 358.3B
84.6%
11 / 38Qwen3 235B A22B Thinking 2507 · 235.1B
84.5%
12 / 41DeepSeek R1 · 684.5B
84.0%
13 / 46DeepSeek R1 0528 · 684.5B
83.4%
14 / 49Qwen3 235B A22B Instruct 2507 · 235.1B
83.0%
15 / 51LongCat Flash Chat · 561.9B
82.7%
16 / 52Seed OSS 36B Instruct · 36.2B
82.7%
17 / 53Qwen3.5 9B · 9.7B
82.5%
18 / 54MiniMax M2 · 228.7B
82.0%
19 / 56GLM 4.5 Air · 110.5B
81.4%
20 / 57DeepSeek v3 0324 · 684.5B
81.3%
21 / 59Kimi K2 Instruct · 1026.4B
81.0%
22 / 60Qwen3 30B A3B Thinking 2507 · 30.5B
80.9%
23 / 61GPT OSS 120B · 116.8B
80.8%
24 / 62Llama 4 Maverick 17B 128E Instruct · 401.6B
80.5%
25 / 65MiniMax M2.5 · 228.7B
80.1%
26 / 69Qwen3.5 4B · 4.7B
79.1%
27 / 74NVIDIA Nemotron 3 Nano 30B A3B BF16 · 31.6B
78.3%
28 / 81Phi 4 Reasoning Plus · 14.7B
76.0%
29 / 82DeepSeek v3 · 684.5B
75.9%
30 / 83MiniMax Text 01 · 456.1B
75.7%
31 / 87Llama 4 Scout 17B 16E Instruct · 108.6B
74.3%
32 / 88Phi 4 Reasoning · 14.7B
74.3%
33 / 89GPT OSS 20B · 20.9B
73.6%
34 / 90Llama 3.1 405B Instruct · 405.9B
73.3%
35 / 95Qwen2.5 72B Instruct · 72.7B
71.6%
36 / 97QwQ 32B Preview · 32.8B
71.0%
37 / 98Phi 4 · 14.7B
70.4%
38 / 100Athene v2 Chat · 72.7B
70.2%
39 / 102Qwen2.5 32B Instruct · 32.8B
69.2%
40 / 104QwQ 32B · 32.8B
69.1%
41 / 107Qwen3 235B A22B · 235.1B
68.2%
42 / 108Mistral Large Instruct 2411 · 122.6B
67.9%
43 / 109Gemma 3 27B IT · 27.4B
67.5%
44 / 114Llama 3.3 70B Instruct · 70.6B
65.9%
45 / 115Mistral Large Instruct 2407 · 122.6B
65.9%
46 / 116DeepSeek V2.5 · 235.7B
65.8%
47 / 117NVIDIA Nemotron 3 Nano 30B A3B Base BF16 · 31.6B
65.1%
48 / 118Seed OSS 36B Base · 36.2B
65.1%
49 / 120Qwen2 72B Instruct · 72.7B
64.4%
50 / 124Qwen2.5 14B · 14.8B
63.7%
51 / 125DeepSeek Coder v2 Instruct · 235.7B
63.6%
52 / 126Higgs Llama 3 70B · 70.6B
63.2%
53 / 129Llama 3.1 70B Instruct · 70.6B
62.8%
54 / 130Llama 3.1 Nemotron 70B Instruct HF · 70.6B
62.8%
55 / 134Qwen3 30B A3B Base · 30.5B
61.7%
56 / 135Llama 3.1 405B · 405.9B
61.6%
57 / 136Gemma 3 12B IT · 12.2B
60.6%
58 / 138Reflection Llama 3.1 70B · 70.6B
60.4%
59 / 141EXAONE 3.5 32B Instruct · 32.0B
58.9%
60 / 143MiMo 7B RL · 7.8B
58.6%
61 / 146Internlm3 8B Instruct · 8.8B
57.6%
62 / 149Gemma 2 27B IT · 27.2B
56.5%
63 / 150Mixtral 8x22B Instruct v0.1 · 140.6B
56.3%
64 / 151Meta Llama 3 70B Instruct · 70.6B
56.2%
65 / 152Phi 3 Medium 4k Instruct · 14.0B
55.7%
66 / 155Qwen3.5 2B · 2.3B
55.3%
67 / 158Phi 4 Mini Instruct · 3.8B
52.8%
68 / 159Meta Llama 3 70B · 70.6B
52.8%
69 / 160Qwen1.5 72B Chat · 72.3B
52.6%
70 / 161Llama 3.1 70B · 70.6B
52.5%
71 / 162Yi 1.5 34B Chat · 34.4B
52.3%
72 / 163Gemma 2 9B IT · 9.2B
52.1%
73 / 164Phi 3 Medium 128k Instruct · 14.0B
51.9%
74 / 166Qwen1.5 110B · 111.2B
49.9%
75 / 167AI21 Jamba Large 1.5 · 398.6B
49.5%
76 / 168Mistral Small Instruct 2409 · 22.2B
48.4%
77 / 169Glm 4 9B Chat · 9.4B
48.0%
78 / 171Phi 3.5 Mini Instruct · 3.8B
47.9%
79 / 172Qwen2 7B Instruct · 7.6B
47.2%
80 / 174EXAONE 3.5 7.8B Instruct · 7.8B
46.2%
81 / 175Yi 1.5 9B Chat · 8.8B
46.0%
82 / 176Phi 3 Mini 4k Instruct · 3.8B
45.7%
83 / 177Aya Expanse 32B · 32.3B
45.4%
84 / 178Gemma 2 9B · 9.2B
45.1%
85 / 179Qwen2.5 7B · 7.6B
45.0%
86 / 180Mistral Nemo Instruct 2407 · 12.2B
44.8%
87 / 181Llama 3.1 8B Instruct · 8.0B
44.3%
88 / 182Nemotron H 8B Base 8K · 8.1B
44.0%
89 / 183Phi 3 Mini 128k Instruct · 3.8B
43.9%
90 / 184Qwen2.5 3B · 3.1B
43.7%
91 / 185Gemma 3 4B IT · 4.3B
43.6%
92 / 187Mixtral 8x7B Instruct v0.1 · 46.7B
43.3%
93 / 188Yi 34B · 34.4B
43.0%
94 / 190Mathstral 7B v0.1 · 7.2B
42.0%
95 / 191MiMo 7B Base · 7.8B
41.9%
96 / 192DeepSeek Coder v2 Lite Instruct · 15.7B
41.6%
97 / 193Granite 3.1 8B Instruct · 8.2B
41.0%
98 / 194Mixtral 8x7B v0.1 · 46.7B
41.0%
99 / 195Meta Llama 3 8B Instruct · 8.0B
41.0%
100 / 197Qwen2 7B · 7.6B
40.7%
101 / 198Mistral Nemo Base 2407 · 12.2B
39.8%
102 / 199WizardLM 2 8x22B · 140.6B
39.2%
103 / 200EXAONE 3.5 2.4B Instruct · 2.4B
39.1%
104 / 201Yi 1.5 6B Chat · 6.1B
38.2%
105 / 202Qwen1.5 14B Chat · 14.2B
38.0%
106 / 203Ministral 8B Instruct 2410 · 8.0B
37.9%
107 / 204C4ai Command R V01 · 35.0B
37.9%
108 / 206Llama 2 70B HF · 69.0B
37.5%
109 / 211Llama 3.1 8B · 8.0B
36.6%
110 / 212Meta Llama 3 8B · 8.0B
35.4%
111 / 214DeepSeek Coder v2 Lite Base · 15.7B
34.4%
112 / 215Aya Expanse 8B · 8.0B
33.7%
113 / 216Gemma 7B · 8.5B
33.7%
114 / 219Zephyr 7B Beta · 7.2B
33.0%
115 / 220Qwen2.5 1.5B · 1.5B
32.1%
116 / 221Granite 3.1 2B Instruct · 2.5B
32.0%
117 / 223Mistral 7B v0.1 · 7.2B
30.9%
118 / 224Mistral 7B Instruct v0.2 · 7.2B
30.8%
119 / 225Mistral 7B v0.2 · 7.2B
30.4%
120 / 226Qwen3.5 0.8B · 873M
29.7%
121 / 227Qwen1.5 7B Chat · 7.7B
29.1%
122 / 228Yi 6B Chat · 6.1B
28.8%
123 / 230Yi 6B · 6.1B
26.5%
124 / 232Mistral 7B Instruct v0.1 · 7.2B
25.8%
125 / 234Llama 2 13B HF · 13.0B
25.3%
126 / 236Llemma 7B · 7B
23.4%
127 / 237Qwen2 1.5B Instruct · 1.5B
22.6%
128 / 238Qwen2 1.5B · 1.5B
22.6%
129 / 239Llama 3.2 3B · 3.2B
22.2%
130 / 242Llama 2 7B HF · 6.7B
20.3%
131 / 243SmolLM2 1.7B · 1.7B
18.3%
132 / 244Qwen2 0.5B Instruct · 494M
15.9%
133 / 245Gemma 2B · 2.5B
15.8%
134 / 246Gemma 2 2B IT · 2.6B
15.6%
135 / 247Qwen2 0.5B · 494M
15.0%
136 / 248Qwen2.5 0.5B · 494M
14.9%
137 / 249Gemma 3 1B IT · 1000M
14.7%
138 / 251Granite 3.1 1B A400m Base · 1.3B
12.3%
139 / 252Llama 3.2 1B · 1.2B
11.9%
140 / 253SmolLM 1.7B · 1.7B
11.9%
141 / 254SmolLM2 360M · 362M
11.4%
142 / 255SmolLM 135M · 135M
11.2%
143 / 256SmolLM 360M · 362M
10.9%
144 / 257SmolLM2 135M · 135M
10.8%

Score vs model size

Which models give the most quality for their size — the ones worth running locally.

1B10B100B1Tmodel size (log scale) →88.0%10.8%Qwen3.5 397B A17B · 403B · 87.8%Kimi K2.5 · 1T · 87.1%GLM 5 · 754B · 86.0%Qwen3.5 35B A3B · 36B · 85.3%DeepSeek V3.2 · 685B · 85.0%DeepSeek V3.1 · 685B · 84.8%GLM 4.5 · 358B · 84.6%Qwen3 235B A22B Thinking 2507 · 235B · 84.5%DeepSeek R1 · 684B · 84.0%DeepSeek R1 0528 · 685B · 83.4%Qwen3 235B A22B Instruct 2507 · 235B · 83.0%Seed OSS 36B Instruct · 36B · 82.7%LongCat Flash Chat · 562B · 82.7%MiniMax M2 · 229B · 82.0%GLM 4.5 Air · 110B · 81.4%DeepSeek v3 0324 · 685B · 81.3%Kimi K2 Instruct · 1T · 81.0%Qwen3 30B A3B Thinking 2507 · 31B · 80.9%GPT OSS 120B · 117B · 80.8%Llama 4 Maverick 17B 128E Instruct · 402B · 80.5%MiniMax M2.5 · 229B · 80.1%NVIDIA Nemotron 3 Nano 30B A3B BF16 · 32B · 78.3%Phi 4 Reasoning Plus · 15B · 76.0%DeepSeek v3 · 685B · 75.9%MiniMax Text 01 · 456B · 75.7%Llama 4 Scout 17B 16E Instruct · 109B · 74.3%Phi 4 Reasoning · 15B · 74.3%GPT OSS 20B · 21B · 73.6%Llama 3.1 405B Instruct · 406B · 73.3%Qwen2.5 72B Instruct · 73B · 71.6%QwQ 32B Preview · 33B · 71.0%Phi 4 · 15B · 70.4%Athene v2 Chat · 73B · 70.2%Qwen2.5 32B Instruct · 33B · 69.2%QwQ 32B · 33B · 69.1%Qwen3 235B A22B · 235B · 68.2%Mistral Large Instruct 2411 · 123B · 67.9%Gemma 3 27B IT · 27B · 67.5%Llama 3.3 70B Instruct · 71B · 65.9%Mistral Large Instruct 2407 · 123B · 65.9%DeepSeek V2.5 · 236B · 65.8%Seed OSS 36B Base · 36B · 65.1%NVIDIA Nemotron 3 Nano 30B A3B Base BF16 · 32B · 65.1%Qwen2 72B Instruct · 73B · 64.4%Qwen2.5 14B · 15B · 63.7%DeepSeek Coder v2 Instruct · 236B · 63.6%Higgs Llama 3 70B · 71B · 63.2%Llama 3.1 70B Instruct · 71B · 62.8%Llama 3.1 Nemotron 70B Instruct HF · 71B · 62.8%Qwen3 30B A3B Base · 31B · 61.7%Llama 3.1 405B · 406B · 61.6%Gemma 3 12B IT · 12B · 60.6%Reflection Llama 3.1 70B · 71B · 60.4%EXAONE 3.5 32B Instruct · 32B · 58.9%MiMo 7B RL · 8B · 58.6%Internlm3 8B Instruct · 9B · 57.6%Gemma 2 27B IT · 27B · 56.5%Mixtral 8x22B Instruct v0.1 · 141B · 56.3%Meta Llama 3 70B Instruct · 71B · 56.2%Phi 3 Medium 4k Instruct · 14B · 55.7%Phi 4 Mini Instruct · 4B · 52.8%Meta Llama 3 70B · 71B · 52.8%Qwen1.5 72B Chat · 72B · 52.6%Llama 3.1 70B · 71B · 52.5%Yi 1.5 34B Chat · 34B · 52.3%Gemma 2 9B IT · 9B · 52.1%Phi 3 Medium 128k Instruct · 14B · 51.9%Qwen1.5 110B · 111B · 49.9%AI21 Jamba Large 1.5 · 399B · 49.5%Mistral Small Instruct 2409 · 22B · 48.4%Glm 4 9B Chat · 9B · 48.0%Phi 3.5 Mini Instruct · 4B · 47.9%Qwen2 7B Instruct · 8B · 47.2%EXAONE 3.5 7.8B Instruct · 8B · 46.2%Yi 1.5 9B Chat · 9B · 46.0%Phi 3 Mini 4k Instruct · 4B · 45.7%Aya Expanse 32B · 32B · 45.4%Gemma 2 9B · 9B · 45.1%Qwen2.5 7B · 8B · 45.0%Mistral Nemo Instruct 2407 · 12B · 44.8%Llama 3.1 8B Instruct · 8B · 44.3%Nemotron H 8B Base 8K · 8B · 44.0%Phi 3 Mini 128k Instruct · 4B · 43.9%Qwen2.5 3B · 3B · 43.7%Gemma 3 4B IT · 4B · 43.6%Mixtral 8x7B Instruct v0.1 · 47B · 43.3%Yi 34B · 34B · 43.0%Mathstral 7B v0.1 · 7B · 42.0%MiMo 7B Base · 8B · 41.9%DeepSeek Coder v2 Lite Instruct · 16B · 41.6%Granite 3.1 8B Instruct · 8B · 41.0%Mixtral 8x7B v0.1 · 47B · 41.0%Meta Llama 3 8B Instruct · 8B · 41.0%Qwen2 7B · 8B · 40.7%Mistral Nemo Base 2407 · 12B · 39.8%WizardLM 2 8x22B · 141B · 39.2%EXAONE 3.5 2.4B Instruct · 2B · 39.1%Yi 1.5 6B Chat · 6B · 38.2%Qwen1.5 14B Chat · 14B · 38.0%Ministral 8B Instruct 2410 · 8B · 37.9%C4ai Command R V01 · 35B · 37.9%Llama 2 70B HF · 69B · 37.5%Llama 3.1 8B · 8B · 36.6%Meta Llama 3 8B · 8B · 35.4%DeepSeek Coder v2 Lite Base · 16B · 34.4%Aya Expanse 8B · 8B · 33.7%Gemma 7B · 9B · 33.7%Zephyr 7B Beta · 7B · 33.0%Granite 3.1 2B Instruct · 3B · 32.0%Mistral 7B v0.1 · 7B · 30.9%Mistral 7B Instruct v0.2 · 7B · 30.8%Mistral 7B v0.2 · 7B · 30.4%Qwen1.5 7B Chat · 8B · 29.1%Yi 6B Chat · 6B · 28.8%Yi 6B · 6B · 26.5%Mistral 7B Instruct v0.1 · 7B · 25.8%Llama 2 13B HF · 13B · 25.3%Llemma 7B · 7B · 23.4%Qwen2 1.5B Instruct · 2B · 22.6%Qwen2 1.5B · 2B · 22.6%Llama 3.2 3B · 3B · 22.2%Llama 2 7B HF · 7B · 20.3%SmolLM2 1.7B · 2B · 18.3%Gemma 2B · 3B · 15.8%Gemma 2 2B IT · 3B · 15.6%Qwen2 0.5B · 494M · 15.0%Qwen2.5 0.5B · 494M · 14.9%Gemma 3 1B IT · 1000M · 14.7%Granite 3.1 1B A400m Base · 1B · 12.3%Llama 3.2 1B · 1B · 11.9%SmolLM 1.7B · 2B · 11.9%SmolLM 360M · 362M · 10.9%SmolLM2 135M · 135M · 10.8%SmolLM 135M · 135M · 11.2%SmolLM 135MSmolLM2 360M · 362M · 11.4%SmolLM2 360MQwen2 0.5B Instruct · 494M · 15.9%Qwen2 0.5B InstructQwen3.5 0.8B · 873M · 29.7%Qwen3.5 0.8BQwen2.5 1.5B · 2B · 32.1%Qwen2.5 1.5BQwen3.5 2B · 2B · 55.3%Qwen3.5 2BQwen3.5 4B · 5B · 79.1%Qwen3.5 4BQwen3.5 9B · 10B · 82.5%Qwen3.5 9BQwen3.5 27B · 28B · 86.1%Qwen3.5 27BQwen3.5 122B A10B · 125B · 86.7%Qwen3.5 122B A10BMiniMax M2.1 · 229B · 88.0%MiniMax M2.1
Each dot is a model. Up = higher score, left = smaller (easier to run locally). The dashed line marks the efficiency frontier — the best score you can get at each size or smaller.
  • SmolLM 135M, 135M, score 11.2% — on the efficiency frontier (best score at its size or smaller).
  • SmolLM2 360M, 362M, score 11.4% — on the efficiency frontier (best score at its size or smaller).
  • Qwen2 0.5B Instruct, 494M, score 15.9% — on the efficiency frontier (best score at its size or smaller).
  • Qwen3.5 0.8B, 873M, score 29.7% — on the efficiency frontier (best score at its size or smaller).
  • Qwen2.5 1.5B, 2B, score 32.1% — on the efficiency frontier (best score at its size or smaller).
  • Qwen3.5 2B, 2B, score 55.3% — on the efficiency frontier (best score at its size or smaller).
  • Qwen3.5 4B, 5B, score 79.1% — on the efficiency frontier (best score at its size or smaller).
  • Qwen3.5 9B, 10B, score 82.5% — on the efficiency frontier (best score at its size or smaller).
  • Qwen3.5 27B, 28B, score 86.1% — on the efficiency frontier (best score at its size or smaller).
  • Qwen3.5 122B A10B, 125B, score 86.7% — on the efficiency frontier (best score at its size or smaller).
  • MiniMax M2.1, 229B, score 88.0% — on the efficiency frontier (best score at its size or smaller).

MMLU-Pro: frequently asked questions

What is the best open LLM on MMLU-Pro?
MiniMax M2.1 is the top open model on MMLU-Pro, scoring 88.0%. Among all models tested — including proprietary ones — it ranks #6. The top model overall is Gemini-3.1-Pro (Google) at 91.2%.
What's the best MMLU-Pro model you can run on a 24 GB GPU?
Qwen3.5 27B is the highest-scoring open model that fits in 24 GB at 4-bit quantization (about 15 GB), scoring 86.1% on MMLU-Pro.
What's the best MMLU-Pro model you can run on a 12 GB GPU?
Qwen3.5 9B is the highest-scoring open model that fits in 12 GB at 4-bit quantization (about 5 GB), scoring 82.5% on MMLU-Pro.
Can open models match proprietary models on MMLU-Pro?
Not quite on MMLU-Pro: the strongest proprietary model (Gemini-3.1-Pro) scores 91.2%, ahead of the best open model (MiniMax M2.1) at 88.0% — but you can run the open one yourself.

Scores aggregated from tigerlab. llmrun does not run this benchmark — see the source for methodology, or the about benchmarks for what it measures.