Reasoning

GPQA Diamond Leaderboard

GPQA Diamond is a set of extremely hard, graduate-level science questions (physics, chemistry, biology) written by domain experts and filtered so that skilled non-experts with web access still fail. It measures genuine reasoning rather than memorization.

Source: epoch83 open models ranked+208 proprietaryData through Sep 2026

Open models ranked on GPQA Diamond

# shows rank among open models / rank overall (including proprietary).

#ModelScore
1 / 16Kimi K3 · 2779.9B
93.1%
2 / 21GLM 5.2 · 753.3B
91.9%
3 / 22DeepSeek V4 Pro 0813 · 1650.5B
91.7%
4 / 26DeepSeek V4 Flash 0731 · 304.2B
91.0%
5 / 27DeepSeek V4 Pro · 1598.8B
90.9%
6 / 28GLM 5.3 · 753.3B
90.9%
7 / 29MiniMax M3 · 427.0B
90.9%
8 / 31Kimi K2.6 · 1026.9B
90.8%
9 / 36GLM 5.3 Flash · 321.3B
90.1%
10 / 37GLM 5.1 · 753.9B
89.9%
11 / 47Inkling Small · 266.0B
88.5%
12 / 51Inkling · 952.4B
88.3%
13 / 55Kimi K2.7 Code · 1026.9B
87.9%
14 / 57GLM 5 · 753.9B
87.8%
15 / 59Kimi K2.5 · 1026.9B
87.6%
16 / 68Qwen3.5 397B A17B · 403.4B
86.4%
17 / 73Qwen3.6 27B · 27.8B
85.9%
18 / 83Qwen3.6 35B A3B · 36.0B
84.9%
19 / 84Kimi K2 Thinking · 1026.4B
84.2%
20 / 93GLM 4.7 · 358.3B
83.3%
21 / 112Qwen3 235B A22B Thinking 2507 · 235.1B
80.0%
22 / 115Qwen3.5 9B · 9.7B
79.0%
23 / 132DeepSeek R1 0528 · 684.5B
76.3%
24 / 137Gemma 4 31B IT · 31.3B
75.8%
25 / 138GPT OSS 120B · 116.8B
75.8%
26 / 151Gemma 4 26B A4B IT · 25.8B
73.2%
27 / 158Seed OSS 36B Instruct · 36.2B
71.5%
28 / 161Qwen3 235B A22B · 235.1B
70.7%
29 / 162Qwen3 30B A3B Thinking 2507 · 30.5B
70.1%
30 / 164DeepSeek R1 · 684.5B
69.2%
31 / 168DeepSeek v3 0324 · 684.5B
67.6%
32 / 171Llama 4 Maverick 17B 128E Instruct · 401.6B
67.0%
33 / 178Qwen3 32B · 32.8B
65.7%
34 / 181QwQ 32B · 32.8B
65.3%
35 / 182DeepSeek R1 Distill Qwen 32B · 32.8B
64.1%
36 / 185Qwen3 14B · 14.8B
63.8%
37 / 188Qwen3 30B A3B · 30.5B
61.7%
38 / 189GPT OSS 20B · 20.9B
60.8%
39 / 190GLM 4.7 Flash · 31.2B
60.5%
40 / 197Qwen3 8B · 8.2B
56.8%
41 / 198DeepSeek v3 · 684.5B
56.5%
42 / 200Magistral Small 2506 · 23.6B
56.1%
43 / 201Phi 4 · 14.7B
56.1%
44 / 202DeepSeek R1 Distill Llama 70B · 70.6B
55.7%
45 / 203Qwen3 30B A3B Instruct 2507 · 30.5B
55.6%
46 / 208Qwen3 4B · 4.0B
52.3%
47 / 209Llama 4 Scout 17B 16E Instruct · 108.6B
51.8%
48 / 211Llama 3.1 405B Instruct · 405.9B
50.9%
49 / 214Qwen2.5 72B Instruct · 72.7B
49.1%
50 / 222Gemma 3 27B IT · 27.4B
47.7%
51 / 225Llama 3.3 70B Instruct · 70.6B
47.4%
52 / 230Llama 3.1 Tulu 3 70B DPO · 70.6B
46.3%
53 / 231Qwen2.5 32B Instruct · 32.8B
46.1%
54 / 233Qwen3 4B Instruct 2507 · 4.0B
45.8%
55 / 235DeepSeek R1 Distill Qwen 14B · 14.8B
44.7%
56 / 236Llama 3.1 70B Instruct · 70.6B
44.2%
57 / 237WizardLM 2 8x22B · 140.6B
43.4%
58 / 242Llama 3.2 90B Vision Instruct · 88.6B
41.0%
59 / 243Qwen2 72B Instruct · 72.7B
40.8%
60 / 245Meta Llama 3 70B Instruct · 70.6B
40.6%
61 / 247Gemma 3 12B IT · 12.2B
39.5%
62 / 250Qwen3 1.7B · 2.0B
38.0%
63 / 252Hermes 2 Theta Llama 3 70B · 70.6B
37.5%
64 / 253Gemma 2 27B IT · 27.2B
36.5%
65 / 256Qwen2.5 7B Instruct · 7.6B
35.5%
66 / 260Eurus 2 7B PRIME · 7.6B
33.9%
67 / 261DeepSeek R1 Distill Qwen 1.5B · 1.8B
33.6%
68 / 265Yi 1.5 34B Chat · 34.4B
32.0%
69 / 266Qwen1.5 32B Chat · 32.5B
30.7%
70 / 268Mixtral 8x7B Instruct v0.1 · 46.7B
30.6%
71 / 271Qwen1.5 72B Chat · 72.3B
28.8%
72 / 272Granite 4.0 Micro · 3.4B
28.3%
73 / 275Gemma 2 9B IT · 9.2B
27.5%
74 / 278Llama 3.1 8B Instruct · 8.0B
27.0%
75 / 279Llama 2 70B Chat HF · 69.0B
26.3%
76 / 280Meta Llama 3 8B Instruct · 8.0B
26.1%
77 / 282Deepseek Llm 67B Chat · 67B
24.6%
78 / 284Llama 3.2 1B Instruct · 1.2B
23.9%
79 / 285Gemma 3 4B IT · 4.3B
23.2%
80 / 286Gemma 3 1B IT · 1000M
20.0%
81 / 287Mistral 7B Instruct v0.3 · 7.2B
15.2%
82 / 288Yi 34B Chat · 34.4B
14.7%
83 / 291DeepSeek R1 0528 Qwen3 8B · 8.2B
9.3%

Score vs model size

Which models give the most quality for their size — the ones worth running locally.

1B10B100B1Tmodel size (log scale) →93.1%9.3%DeepSeek V4 Pro 0813 · 1.7T · 91.7%DeepSeek V4 Pro · 1.6T · 90.9%MiniMax M3 · 427B · 90.9%GLM 5.3 · 753B · 90.9%Kimi K2.6 · 1T · 90.8%GLM 5.3 Flash · 321B · 90.1%GLM 5.1 · 754B · 89.9%Inkling · 952B · 88.3%Kimi K2.7 Code · 1T · 87.9%GLM 5 · 754B · 87.8%Kimi K2.5 · 1T · 87.6%Qwen3.5 397B A17B · 403B · 86.4%Qwen3.6 35B A3B · 36B · 84.9%Kimi K2 Thinking · 1T · 84.2%GLM 4.7 · 358B · 83.3%Qwen3 235B A22B Thinking 2507 · 235B · 80.0%DeepSeek R1 0528 · 685B · 76.3%Gemma 4 31B IT · 31B · 75.8%GPT OSS 120B · 117B · 75.8%Gemma 4 26B A4B IT · 26B · 73.2%Seed OSS 36B Instruct · 36B · 71.5%Qwen3 235B A22B · 235B · 70.7%Qwen3 30B A3B Thinking 2507 · 31B · 70.1%DeepSeek R1 · 685B · 69.2%DeepSeek v3 0324 · 685B · 67.6%Llama 4 Maverick 17B 128E Instruct · 402B · 67.0%Qwen3 32B · 33B · 65.7%QwQ 32B · 33B · 65.3%DeepSeek R1 Distill Qwen 32B · 33B · 64.1%Qwen3 14B · 15B · 63.8%Qwen3 30B A3B · 31B · 61.7%GPT OSS 20B · 21B · 60.8%GLM 4.7 Flash · 31B · 60.5%DeepSeek v3 · 685B · 56.5%Phi 4 · 15B · 56.1%Magistral Small 2506 · 24B · 56.1%DeepSeek R1 Distill Llama 70B · 71B · 55.7%Qwen3 30B A3B Instruct 2507 · 31B · 55.6%Llama 4 Scout 17B 16E Instruct · 109B · 51.8%Llama 3.1 405B Instruct · 406B · 50.9%Qwen2.5 72B Instruct · 73B · 49.1%Gemma 3 27B IT · 27B · 47.7%Llama 3.3 70B Instruct · 71B · 47.4%Llama 3.1 Tulu 3 70B DPO · 71B · 46.3%Qwen2.5 32B Instruct · 33B · 46.1%Qwen3 4B Instruct 2507 · 4B · 45.8%DeepSeek R1 Distill Qwen 14B · 15B · 44.7%Llama 3.1 70B Instruct · 71B · 44.2%WizardLM 2 8x22B · 141B · 43.4%Llama 3.2 90B Vision Instruct · 89B · 41.0%Qwen2 72B Instruct · 73B · 40.8%Meta Llama 3 70B Instruct · 71B · 40.6%Gemma 3 12B IT · 12B · 39.5%Hermes 2 Theta Llama 3 70B · 71B · 37.5%Gemma 2 27B IT · 27B · 36.5%Qwen2.5 7B Instruct · 8B · 35.5%Eurus 2 7B PRIME · 8B · 33.9%Yi 1.5 34B Chat · 34B · 32.0%Qwen1.5 32B Chat · 33B · 30.7%Mixtral 8x7B Instruct v0.1 · 47B · 30.6%Qwen1.5 72B Chat · 72B · 28.8%Granite 4.0 Micro · 3B · 28.3%Gemma 2 9B IT · 9B · 27.5%Llama 3.1 8B Instruct · 8B · 27.0%Llama 2 70B Chat HF · 69B · 26.3%Meta Llama 3 8B Instruct · 8B · 26.1%Deepseek Llm 67B Chat · 67B · 24.6%Gemma 3 4B IT · 4B · 23.2%Mistral 7B Instruct v0.3 · 7B · 15.2%Yi 34B Chat · 34B · 14.7%DeepSeek R1 0528 Qwen3 8B · 8B · 9.3%Gemma 3 1B IT · 1000M · 20.0%Llama 3.2 1B Instruct · 1B · 23.9%Llama 3.2 1B InstructDeepSeek R1 Distill Qwen 1.5B · 2B · 33.6%DeepSeek R1 Distill Q…Qwen3 1.7B · 2B · 38.0%Qwen3 1.7BQwen3 4B · 4B · 52.3%Qwen3 4BQwen3 8B · 8B · 56.8%Qwen3 8BQwen3.5 9B · 10B · 79.0%Qwen3.5 9BQwen3.6 27B · 28B · 85.9%Qwen3.6 27BInkling Small · 266B · 88.5%Inkling SmallDeepSeek V4 Flash 0731 · 304B · 91.0%DeepSeek V4 Flash 0731GLM 5.2 · 753B · 91.9%GLM 5.2Kimi K3 · 2.8T · 93.1%Kimi K3
Each dot is a model. Up = higher score, left = smaller (easier to run locally). The dashed line marks the efficiency frontier — the best score you can get at each size or smaller.
  • Gemma 3 1B IT, 1000M, score 20.0% — on the efficiency frontier (best score at its size or smaller).
  • Llama 3.2 1B Instruct, 1B, score 23.9% — on the efficiency frontier (best score at its size or smaller).
  • DeepSeek R1 Distill Qwen 1.5B, 2B, score 33.6% — on the efficiency frontier (best score at its size or smaller).
  • Qwen3 1.7B, 2B, score 38.0% — on the efficiency frontier (best score at its size or smaller).
  • Qwen3 4B, 4B, score 52.3% — on the efficiency frontier (best score at its size or smaller).
  • Qwen3 8B, 8B, score 56.8% — on the efficiency frontier (best score at its size or smaller).
  • Qwen3.5 9B, 10B, score 79.0% — on the efficiency frontier (best score at its size or smaller).
  • Qwen3.6 27B, 28B, score 85.9% — on the efficiency frontier (best score at its size or smaller).
  • Inkling Small, 266B, score 88.5% — on the efficiency frontier (best score at its size or smaller).
  • DeepSeek V4 Flash 0731, 304B, score 91.0% — on the efficiency frontier (best score at its size or smaller).
  • GLM 5.2, 753B, score 91.9% — on the efficiency frontier (best score at its size or smaller).
  • Kimi K3, 2.8T, score 93.1% — on the efficiency frontier (best score at its size or smaller).

GPQA Diamond: frequently asked questions

What is the best open LLM on GPQA Diamond?
Kimi K3 is the top open model on GPQA Diamond, scoring 93.1%. Among all models tested — including proprietary ones — it ranks #16. The top model overall is GPT 6 Astra Max (OpenAI) at 95.8%.
What's the best GPQA Diamond model you can run on a 24 GB GPU?
Qwen3.6 27B is the highest-scoring open model that fits in 24 GB at 4-bit quantization (about 15 GB), scoring 85.9% on GPQA Diamond.
What's the best GPQA Diamond model you can run on a 12 GB GPU?
Qwen3.5 9B is the highest-scoring open model that fits in 12 GB at 4-bit quantization (about 5 GB), scoring 79.0% on GPQA Diamond.
Can open models match proprietary models on GPQA Diamond?
Not quite on GPQA Diamond: the strongest proprietary model (GPT 6 Astra Max) scores 95.8%, ahead of the best open model (Kimi K3) at 93.1% — but you can run the open one yourself.

Scores aggregated from epoch. llmrun does not run this benchmark — see the source for methodology, or the about benchmarks for what it measures.