Reasoning

GPQA Diamond Leaderboard

GPQA Diamond is a set of extremely hard, graduate-level science questions (physics, chemistry, biology) written by domain experts and filtered so that skilled non-experts with web access still fail. It measures genuine reasoning rather than memorization.

Source: epoch46 open models ranked+136 proprietaryData through Jul 2026

All models ranked on GPQA Diamond

Proprietary / closed models are shown dimmed — you can't run them locally, but they show where the open field stands.

#ModelScore
1GPT 5.4 Pro (Mar 05, 2026, xhigh) · proprietary
94.6%
2Gemini 3.1 Pro Preview · proprietary
94.1%
3GPT 5.5 Pre Release (xhigh) · proprietary
94.0%
4GPT 5.5 Pro Pre Release (xhigh) · proprietary
93.9%
5GPT 5.6 Sol Max · proprietary
93.5%
6Grok 4.5 (high) · proprietary
93.4%
7GPT 5.6 Terra Max · proprietary
93.3%
8GPT 5.4 (Mar 05, 2026, xhigh) · proprietary
93.3%
9Kimi K3 Max · proprietary
93.1%
10Gemini 3.5 Flash (high) · proprietary
92.8%
11Gemini 3 Pro Preview · proprietary
92.6%
12GLM 5.2 Max · proprietary
91.9%
13GPT 5.6 Luna Max · proprietary
91.6%
14Qwen3.7 Max · proprietary
91.6%
15GPT 5.2 (Dec 11, 2025, xhigh) · proprietary
91.4%
16Claude Opus 4.8 Max · proprietary
91.0%
17Kimi K2.6 · 1058.6B
90.8%
18GPT 5.5 (low) · proprietary
90.7%
19Claude Opus 4.6 (32K) · proprietary
90.5%
20Claude Sonnet 5 (xhigh) · proprietary
90.5%
21Claude Opus 4.7 (xhigh) · proprietary
90.1%
22Muse Spark · proprietary
89.8%
23DeepSeek V4 Pro · 861.6B
89.6%
24Kimi K2.7 Code · 1058.6B
89.5%
25Grok 4.20 0309 Reasoning · proprietary
89.3%
26Qwen3.6 Max Preview · proprietary
89.1%
27Grok 4.3 (high) · proprietary
88.8%
28Claude Opus 4.6 (64K) · proprietary
88.8%
29GPT 5.2 (Dec 11, 2025, high) · proprietary
88.2%
30GPT 5.2 (Dec 11, 2025, medium) · proprietary
87.9%
31GLM 5 · 753.9B
87.8%
32GPT 5.1 (Nov 13, 2025, high) · proprietary
87.6%
33Kimi K2.5 · 1058.6B
87.6%
34Claude Sonnet 4.6 (32K) · proprietary
87.4%
35Qwen3.6 Plus · proprietary
87.4%
36Grok 4 (Jul 09) · proprietary
87.0%
37GPT 5 (Aug 07, 2025, high) · proprietary
86.2%
38Claude Opus 4.5 (Nov 01, 2025, 32K) · proprietary
86.1%
39Claude Opus 4.5 (Nov 01, 2025, 16K) · proprietary
85.5%
40GLM 5.1 · 753.9B
85.5%
41GPT 5 (Aug 07, 2025, medium) · proprietary
85.4%
42Gemini 2.5 Pro · proprietary
85.3%
43GPT 5.1 (Nov 13, 2025, medium) · proprietary
85.0%
44Gemini 2.5 Pro Preview (Jun 05) · proprietary
84.9%
45Qwen3.6 Flash · proprietary
84.4%
46Kimi K2 Thinking · 1058.1B
84.2%
47Qwen3.5 Plus · proprietary
84.2%
48Gemini 2.5 Pro Exp (Mar 25) · proprietary
83.8%
49Qwen3.5 Flash · proprietary
83.8%
50GPT 5.4 Mini (Mar 17, 2026, high) · proprietary
83.6%
51DeepSeek Reasoner · proprietary
83.4%
52GLM 4.7 · 358.3B
83.3%
53Gemini 3 Flash Preview · proprietary
83.2%
54GPT 5.2 (Dec 11, 2025, low) · proprietary
82.7%
55Claude Sonnet 4.5 (Sep 29, 2025, 59K) · proprietary
82.3%
56O3 (Apr 16, 2025, high) · proprietary
81.8%
57Claude Sonnet 4.5 (Sep 29, 2025, 32K) · proprietary
81.7%
58Claude Opus 4.5 (Nov 01, 2025) · proprietary
80.7%
59Qwen3 235B A22B Thinking 2507 · 235.1B
80.0%
60O4 Mini (Apr 16, 2025, high) · proprietary
79.6%
61Claude Sonnet 4.5 (Sep 29, 2025, 16K) · proprietary
78.8%
62Claude 3.7 Sonnet (Feb 19, 2025, 64K) · proprietary
78.5%
63GPT 5.4 Nano (Mar 17, 2026, high) · proprietary
78.5%
64Claude Sonnet 4 (May 14, 2025, 32K) · proprietary
78.3%
65Claude Sonnet 4 (May 14, 2025, 59K) · proprietary
77.8%
66Claude Opus 4.1 (Aug 05, 2025, 16K) · proprietary
77.3%
67O3 Mini (Jan 31, 2025, high) · proprietary
77.0%
68Claude 3.7 Sonnet (Feb 19, 2025, 16K) · proprietary
76.8%
69Claude 3.7 Sonnet (Feb 19, 2025, 32K) · proprietary
76.8%
70Claude Opus 4.1 (Aug 05, 2025, 27K) · proprietary
76.8%
71O1 (Dec 17, 2024, high) · proprietary
76.8%
72DeepSeek R1 0528 · 684.5B
76.3%
73Claude Opus 4 (May 14, 2025, 16K) · proprietary
76.3%
74Grok 3 Mini Beta (low) · proprietary
76.3%
75Claude Sonnet 4 (May 14, 2025, 16K) · proprietary
75.8%
76GPT OSS 120B · 120.4B
75.8%
77O1 (Dec 17, 2024, medium) · proprietary
75.8%
78GPT 5 Mini (Aug 07, 2025, high) · proprietary
75.0%
79Grok 3 Mini Beta (high) · proprietary
74.6%
80O3 Mini (Jan 31, 2025, medium) · proprietary
74.3%
81Claude Sonnet 4.5 (Sep 29, 2025) · proprietary
73.7%
82Claude Opus 4.1 (Aug 05, 2025) · proprietary
73.2%
83Qwen3 Max (Sep 23, 2025) · proprietary
72.6%
84GPT 5 Mini (Aug 07, 2025, medium) · proprietary
71.7%
85Claude Haiku 4.5 (Oct 01, 2025, 32K) · proprietary
71.2%
86Qwen3 235B A22B · 235.1B
70.7%
87GPT 5 Nano (Aug 07, 2025, high) · proprietary
69.4%
88DeepSeek R1 · 684.5B
69.2%
89Claude Opus 4 (May 14, 2025) · proprietary
69.2%
90GPT 4.5 Preview (Feb 27, 2025) · proprietary
68.7%
91DeepSeek v3 0324 · 684.5B
67.6%
92Grok 3 Beta · proprietary
67.6%
93GPT 5 Nano (Aug 07, 2025, medium) · proprietary
67.4%
94Llama 4 Maverick 17B 128E Instruct · 401.6B
67.0%
95GPT 4.1 (Apr 14, 2025) · proprietary
66.9%
96Claude Sonnet 4 (May 14, 2025) · proprietary
66.7%
97Gemini 2.5 Pro Preview (May 06) · proprietary
66.7%
98Claude 3.7 Sonnet (Feb 19, 2025) · proprietary
66.0%
99GPT 4.1 Mini (Apr 14, 2025) · proprietary
65.8%
100Gemini 2.0 Pro Exp (Feb 05) · proprietary
65.7%
101QwQ Plus · proprietary
65.4%
102Gemini 2.0 Flash 001 · proprietary
64.1%
103O1 Mini (Sep 12, 2024, high) · proprietary
62.4%
104Claude Haiku 4.5 (Oct 01, 2025) · proprietary
60.5%
105Mistral Medium 2505 · proprietary
59.5%
106O1 Mini (Sep 12, 2024, medium) · proprietary
59.5%
107Gemini 1.5 Pro 002 · proprietary
57.2%
108Gemini 2.0 Flash Thinking Exp (Jan 21) · proprietary
57.1%
109DeepSeek v3 · 684.5B
56.5%
110Qwen Max (Jan 25, 2025) · proprietary
56.1%
111Phi 4 · 14.7B
56.1%
112DeepSeek R1 Distill Llama 70B · 70.6B
55.7%
113Claude 3.5 Sonnet (Oct 22, 2024) · proprietary
55.3%
114Claude 3.5 Sonnet (Jun 20, 2024) · proprietary
54.0%
115Grok 2 (Dec 12) · proprietary
53.8%
116Llama 4 Scout 17B 16E Instruct · 108.6B
51.8%
117Mistral Large 2411 · proprietary
51.3%
118Llama 3.1 405B Instruct · 405.9B
50.9%
119O1 Preview (Sep 12, 2024) · proprietary
50.3%
120GPT 4o (Aug 06, 2024) · proprietary
49.2%
121Qwen2.5 72B Instruct · 72.7B
49.1%
122Mistral Large 2407 · proprietary
49.0%
123GPT 4.1 Nano (Apr 14, 2025) · proprietary
48.9%
124GPT 4o (May 13, 2024) · proprietary
48.9%
125Gemma 3 27B IT · 27.4B
48.9%
126Magistral Small 2506 · 23.6B
48.4%
127Qwen Plus (Jan 25, 2025) · proprietary
48.1%
128GPT 4o (Nov 20, 2024) · proprietary
47.9%
129Mistral Small 2503 · proprietary
47.5%
130Llama 3.3 70B Instruct · 70.6B
47.4%
131Gemini 1.5 Flash 002 · proprietary
47.3%
132Claude 3 Opus (Feb 29, 2024) · proprietary
47.2%
133GPT 4 Turbo (Apr 09, 2024) · proprietary
46.6%
134Llama 3.1 Tulu 3 70B DPO · 70.6B
46.3%
135Qwen2.5 32B Instruct · 32.8B
46.1%
136Gemini 1.5 Pro 001 · proprietary
45.9%
137Mistral Small 2501 · proprietary
45.3%
138DeepSeek R1 Distill Qwen 14B · 14.8B
44.7%
139Llama 3.1 70B Instruct · 70.6B
44.2%
140WizardLM 2 8x22B · 140.6B
43.4%
141GPT 4 1106 Preview · proprietary
42.4%
142GPT 4 0125 Preview · proprietary
42.3%
143Qwen Turbo (Nov 01, 2024) · proprietary
41.8%
144Llama 3.2 90B Vision Instruct · 88.6B
41.0%
145Qwen2 72B Instruct · 72.7B
40.8%
146Claude 3 Sonnet (Feb 29, 2024) · proprietary
40.6%
147Meta Llama 3 70B Instruct · 70.6B
40.6%
148Gemini 1.5 Flash 001 · proprietary
40.4%
149Mistral Large 2402 · proprietary
38.8%
150Claude 3.5 Haiku (Oct 22, 2024) · proprietary
38.1%
151GPT 4o Mini (Jul 18, 2024) · proprietary
37.7%
152Hermes 2 Theta Llama 3 70B · 70.6B
37.5%
153Gemma 2 27B IT · 27.2B
36.5%
154Claude 3 Haiku (Mar 07, 2024) · proprietary
36.3%
155GPT 4 (Mar 14) · proprietary
35.7%
156Claude 2.0 · proprietary
34.7%
157Open Mixtral 8x22b · proprietary
34.1%
158Gemini 1.0 Pro 001 · proprietary
34.0%
159Eurus 2 7B PRIME · 7.6B
33.9%
160Claude 2.1 · proprietary
33.0%
161Gemini 1.5 Flash 8B 001 · proprietary
33.0%
162Dbrx Instruct · proprietary
32.9%
163Yi 1.5 34B Chat · 34.4B
32.0%
164Qwen1.5 32B Chat · 32.5B
30.7%
165GPT 4 (Jun 13) · proprietary
30.6%
166Mixtral 8x7B Instruct v0.1 · 46.7B
30.6%
167Open Mistral Nemo 2407 · proprietary
29.9%
168Open Mixtral 8x7b · proprietary
29.8%
169Qwen1.5 72B Chat · 72.3B
28.8%
170GPT 3.5 Turbo (Nov 06) · proprietary
28.0%
171Phi 3 Medium 128K Instruct · proprietary
27.6%
172Gemma 2 9B IT · 9.2B
27.5%
173GPT 3.5 Turbo (Jan 25) · proprietary
27.2%
174Ministral 8B 2410 · proprietary
27.2%
175Llama 2 70B Chat HF · 69.0B
26.3%
176Meta Llama 3 8B Instruct · 8.0B
26.1%
177Llama 3.1 8B Instruct · 8.0B
25.9%
178Ministral 3B 2410 · proprietary
25.3%
179Deepseek Llm 67B Chat · 67B
24.6%
180Mistral 7B Instruct v0.3 · 7.2B
15.2%
181Yi 34B Chat · 34.4B
14.7%
182Open Mistral 7B · proprietary
13.2%

Score vs model size

Which models give the most quality for their size — the ones worth running locally.

10B100B1Tmodel size (log scale) →90.8%14.7%Kimi K2.7 Code · 1.1T · 89.5%Kimi K2.5 · 1.1T · 87.6%GLM 5.1 · 754B · 85.5%Kimi K2 Thinking · 1.1T · 84.2%DeepSeek R1 0528 · 685B · 76.3%Qwen3 235B A22B · 235B · 70.7%DeepSeek R1 · 685B · 69.2%DeepSeek v3 0324 · 685B · 67.6%Llama 4 Maverick 17B 128E Instruct · 402B · 67.0%DeepSeek v3 · 685B · 56.5%DeepSeek R1 Distill Llama 70B · 71B · 55.7%Llama 4 Scout 17B 16E Instruct · 109B · 51.8%Llama 3.1 405B Instruct · 406B · 50.9%Qwen2.5 72B Instruct · 73B · 49.1%Gemma 3 27B IT · 27B · 48.9%Magistral Small 2506 · 24B · 48.4%Llama 3.3 70B Instruct · 71B · 47.4%Llama 3.1 Tulu 3 70B DPO · 71B · 46.3%Qwen2.5 32B Instruct · 33B · 46.1%DeepSeek R1 Distill Qwen 14B · 15B · 44.7%Llama 3.1 70B Instruct · 71B · 44.2%WizardLM 2 8x22B · 141B · 43.4%Llama 3.2 90B Vision Instruct · 89B · 41.0%Qwen2 72B Instruct · 73B · 40.8%Meta Llama 3 70B Instruct · 71B · 40.6%Hermes 2 Theta Llama 3 70B · 71B · 37.5%Gemma 2 27B IT · 27B · 36.5%Yi 1.5 34B Chat · 34B · 32.0%Qwen1.5 32B Chat · 33B · 30.7%Mixtral 8x7B Instruct v0.1 · 47B · 30.6%Qwen1.5 72B Chat · 72B · 28.8%Gemma 2 9B IT · 9B · 27.5%Llama 2 70B Chat HF · 69B · 26.3%Meta Llama 3 8B Instruct · 8B · 26.1%Llama 3.1 8B Instruct · 8B · 25.9%Deepseek Llm 67B Chat · 67B · 24.6%Yi 34B Chat · 34B · 14.7%Mistral 7B Instruct v0.3 · 7B · 15.2%Mistral 7B Instruct v…Eurus 2 7B PRIME · 8B · 33.9%Eurus 2 7B PRIMEPhi 4 · 15B · 56.1%Phi 4GPT OSS 120B · 120B · 75.8%GPT OSS 120BQwen3 235B A22B Thinking 2507 · 235B · 80.0%Qwen3 235B A22B Think…GLM 4.7 · 358B · 83.3%GLM 4.7GLM 5 · 754B · 87.8%DeepSeek V4 Pro · 862B · 89.6%DeepSeek V4 ProKimi K2.6 · 1.1T · 90.8%Kimi K2.6
Each dot is a model. Up = higher score, left = smaller (easier to run locally). The dashed line marks the efficiency frontier — the best score you can get at each size or smaller.
  • Mistral 7B Instruct v0.3, 7B, score 15.2% — on the efficiency frontier (best score at its size or smaller).
  • Eurus 2 7B PRIME, 8B, score 33.9% — on the efficiency frontier (best score at its size or smaller).
  • Phi 4, 15B, score 56.1% — on the efficiency frontier (best score at its size or smaller).
  • GPT OSS 120B, 120B, score 75.8% — on the efficiency frontier (best score at its size or smaller).
  • Qwen3 235B A22B Thinking 2507, 235B, score 80.0% — on the efficiency frontier (best score at its size or smaller).
  • GLM 4.7, 358B, score 83.3% — on the efficiency frontier (best score at its size or smaller).
  • GLM 5, 754B, score 87.8% — on the efficiency frontier (best score at its size or smaller).
  • DeepSeek V4 Pro, 862B, score 89.6% — on the efficiency frontier (best score at its size or smaller).
  • Kimi K2.6, 1.1T, score 90.8% — on the efficiency frontier (best score at its size or smaller).

GPQA Diamond: frequently asked questions

What is the best open LLM on GPQA Diamond?
Kimi K2.6 is the top open model on GPQA Diamond, scoring 90.8%. Among all models tested — including proprietary ones — it ranks #17. The top model overall is GPT 5.4 Pro (Mar 05, 2026, xhigh) (OpenAI) at 94.6%.
What's the best GPQA Diamond model you can run on a 24 GB GPU?
Phi 4 is the highest-scoring open model that fits in 24 GB at 4-bit quantization (about 8 GB), scoring 56.1% on GPQA Diamond.
What's the best GPQA Diamond model you can run on a 12 GB GPU?
Phi 4 is the highest-scoring open model that fits in 12 GB at 4-bit quantization (about 8 GB), scoring 56.1% on GPQA Diamond.
Can open models match proprietary models on GPQA Diamond?
Not quite on GPQA Diamond: the strongest proprietary model (GPT 5.4 Pro (Mar 05, 2026, xhigh)) scores 94.6%, ahead of the best open model (Kimi K2.6) at 90.8% — but you can run the open one yourself.

Scores aggregated from epoch. llmrun does not run this benchmark — see the source for methodology, or the about benchmarks for what it measures.