Reasoning

GPQA Diamond Leaderboard

GPQA Diamond is a set of extremely hard, graduate-level science questions (physics, chemistry, biology) written by domain experts and filtered so that skilled non-experts with web access still fail. It measures genuine reasoning rather than memorization.

Source: epoch46 open models ranked+136 proprietaryData through Jul 2026

Open models ranked on GPQA Diamond

# shows rank among open models / rank overall (including proprietary).

#ModelScore
1 / 17Kimi K2.6 · 1058.6B
90.8%
2 / 23DeepSeek V4 Pro · 861.6B
89.6%
3 / 24Kimi K2.7 Code · 1058.6B
89.5%
4 / 31GLM 5 · 753.9B
87.8%
5 / 33Kimi K2.5 · 1058.6B
87.6%
6 / 40GLM 5.1 · 753.9B
85.5%
7 / 46Kimi K2 Thinking · 1058.1B
84.2%
8 / 52GLM 4.7 · 358.3B
83.3%
9 / 59Qwen3 235B A22B Thinking 2507 · 235.1B
80.0%
10 / 72DeepSeek R1 0528 · 684.5B
76.3%
11 / 76GPT OSS 120B · 120.4B
75.8%
12 / 86Qwen3 235B A22B · 235.1B
70.7%
13 / 88DeepSeek R1 · 684.5B
69.2%
14 / 91DeepSeek v3 0324 · 684.5B
67.6%
15 / 94Llama 4 Maverick 17B 128E Instruct · 401.6B
67.0%
16 / 109DeepSeek v3 · 684.5B
56.5%
17 / 111Phi 4 · 14.7B
56.1%
18 / 112DeepSeek R1 Distill Llama 70B · 70.6B
55.7%
19 / 116Llama 4 Scout 17B 16E Instruct · 108.6B
51.8%
20 / 118Llama 3.1 405B Instruct · 405.9B
50.9%
21 / 121Qwen2.5 72B Instruct · 72.7B
49.1%
22 / 125Gemma 3 27B IT · 27.4B
48.9%
23 / 126Magistral Small 2506 · 23.6B
48.4%
24 / 130Llama 3.3 70B Instruct · 70.6B
47.4%
25 / 134Llama 3.1 Tulu 3 70B DPO · 70.6B
46.3%
26 / 135Qwen2.5 32B Instruct · 32.8B
46.1%
27 / 138DeepSeek R1 Distill Qwen 14B · 14.8B
44.7%
28 / 139Llama 3.1 70B Instruct · 70.6B
44.2%
29 / 140WizardLM 2 8x22B · 140.6B
43.4%
30 / 144Llama 3.2 90B Vision Instruct · 88.6B
41.0%
31 / 145Qwen2 72B Instruct · 72.7B
40.8%
32 / 147Meta Llama 3 70B Instruct · 70.6B
40.6%
33 / 152Hermes 2 Theta Llama 3 70B · 70.6B
37.5%
34 / 153Gemma 2 27B IT · 27.2B
36.5%
35 / 159Eurus 2 7B PRIME · 7.6B
33.9%
36 / 163Yi 1.5 34B Chat · 34.4B
32.0%
37 / 164Qwen1.5 32B Chat · 32.5B
30.7%
38 / 166Mixtral 8x7B Instruct v0.1 · 46.7B
30.6%
39 / 169Qwen1.5 72B Chat · 72.3B
28.8%
40 / 172Gemma 2 9B IT · 9.2B
27.5%
41 / 175Llama 2 70B Chat HF · 69.0B
26.3%
42 / 176Meta Llama 3 8B Instruct · 8.0B
26.1%
43 / 177Llama 3.1 8B Instruct · 8.0B
25.9%
44 / 179Deepseek Llm 67B Chat · 67B
24.6%
45 / 180Mistral 7B Instruct v0.3 · 7.2B
15.2%
46 / 181Yi 34B Chat · 34.4B
14.7%

Score vs model size

Which models give the most quality for their size — the ones worth running locally.

10B100B1Tmodel size (log scale) →90.8%14.7%Kimi K2.7 Code · 1.1T · 89.5%Kimi K2.5 · 1.1T · 87.6%GLM 5.1 · 754B · 85.5%Kimi K2 Thinking · 1.1T · 84.2%DeepSeek R1 0528 · 685B · 76.3%Qwen3 235B A22B · 235B · 70.7%DeepSeek R1 · 685B · 69.2%DeepSeek v3 0324 · 685B · 67.6%Llama 4 Maverick 17B 128E Instruct · 402B · 67.0%DeepSeek v3 · 685B · 56.5%DeepSeek R1 Distill Llama 70B · 71B · 55.7%Llama 4 Scout 17B 16E Instruct · 109B · 51.8%Llama 3.1 405B Instruct · 406B · 50.9%Qwen2.5 72B Instruct · 73B · 49.1%Gemma 3 27B IT · 27B · 48.9%Magistral Small 2506 · 24B · 48.4%Llama 3.3 70B Instruct · 71B · 47.4%Llama 3.1 Tulu 3 70B DPO · 71B · 46.3%Qwen2.5 32B Instruct · 33B · 46.1%DeepSeek R1 Distill Qwen 14B · 15B · 44.7%Llama 3.1 70B Instruct · 71B · 44.2%WizardLM 2 8x22B · 141B · 43.4%Llama 3.2 90B Vision Instruct · 89B · 41.0%Qwen2 72B Instruct · 73B · 40.8%Meta Llama 3 70B Instruct · 71B · 40.6%Hermes 2 Theta Llama 3 70B · 71B · 37.5%Gemma 2 27B IT · 27B · 36.5%Yi 1.5 34B Chat · 34B · 32.0%Qwen1.5 32B Chat · 33B · 30.7%Mixtral 8x7B Instruct v0.1 · 47B · 30.6%Qwen1.5 72B Chat · 72B · 28.8%Gemma 2 9B IT · 9B · 27.5%Llama 2 70B Chat HF · 69B · 26.3%Meta Llama 3 8B Instruct · 8B · 26.1%Llama 3.1 8B Instruct · 8B · 25.9%Deepseek Llm 67B Chat · 67B · 24.6%Yi 34B Chat · 34B · 14.7%Mistral 7B Instruct v0.3 · 7B · 15.2%Mistral 7B Instruct v…Eurus 2 7B PRIME · 8B · 33.9%Eurus 2 7B PRIMEPhi 4 · 15B · 56.1%Phi 4GPT OSS 120B · 120B · 75.8%GPT OSS 120BQwen3 235B A22B Thinking 2507 · 235B · 80.0%Qwen3 235B A22B Think…GLM 4.7 · 358B · 83.3%GLM 4.7GLM 5 · 754B · 87.8%DeepSeek V4 Pro · 862B · 89.6%DeepSeek V4 ProKimi K2.6 · 1.1T · 90.8%Kimi K2.6
Each dot is a model. Up = higher score, left = smaller (easier to run locally). The dashed line marks the efficiency frontier — the best score you can get at each size or smaller.
  • Mistral 7B Instruct v0.3, 7B, score 15.2% — on the efficiency frontier (best score at its size or smaller).
  • Eurus 2 7B PRIME, 8B, score 33.9% — on the efficiency frontier (best score at its size or smaller).
  • Phi 4, 15B, score 56.1% — on the efficiency frontier (best score at its size or smaller).
  • GPT OSS 120B, 120B, score 75.8% — on the efficiency frontier (best score at its size or smaller).
  • Qwen3 235B A22B Thinking 2507, 235B, score 80.0% — on the efficiency frontier (best score at its size or smaller).
  • GLM 4.7, 358B, score 83.3% — on the efficiency frontier (best score at its size or smaller).
  • GLM 5, 754B, score 87.8% — on the efficiency frontier (best score at its size or smaller).
  • DeepSeek V4 Pro, 862B, score 89.6% — on the efficiency frontier (best score at its size or smaller).
  • Kimi K2.6, 1.1T, score 90.8% — on the efficiency frontier (best score at its size or smaller).

GPQA Diamond: frequently asked questions

What is the best open LLM on GPQA Diamond?
Kimi K2.6 is the top open model on GPQA Diamond, scoring 90.8%. Among all models tested — including proprietary ones — it ranks #17. The top model overall is GPT 5.4 Pro (Mar 05, 2026, xhigh) (OpenAI) at 94.6%.
What's the best GPQA Diamond model you can run on a 24 GB GPU?
Phi 4 is the highest-scoring open model that fits in 24 GB at 4-bit quantization (about 8 GB), scoring 56.1% on GPQA Diamond.
What's the best GPQA Diamond model you can run on a 12 GB GPU?
Phi 4 is the highest-scoring open model that fits in 12 GB at 4-bit quantization (about 8 GB), scoring 56.1% on GPQA Diamond.
Can open models match proprietary models on GPQA Diamond?
Not quite on GPQA Diamond: the strongest proprietary model (GPT 5.4 Pro (Mar 05, 2026, xhigh)) scores 94.6%, ahead of the best open model (Kimi K2.6) at 90.8% — but you can run the open one yourself.

Scores aggregated from epoch. llmrun does not run this benchmark — see the source for methodology, or the about benchmarks for what it measures.