Reasoning
GPQA Diamond Leaderboard
GPQA Diamond is a set of extremely hard, graduate-level science questions (physics, chemistry, biology) written by domain experts and filtered so that skilled non-experts with web access still fail. It measures genuine reasoning rather than memorization.
Source: epoch83 open models ranked+208 proprietaryData through Sep 2026
Open models ranked on GPQA Diamond
# shows rank among open models / rank overall (including proprietary).
| # | Model | Score |
|---|---|---|
| 1 / 16 | Kimi K3 · 2779.9B | 93.1% |
| 2 / 21 | GLM 5.2 · 753.3B | 91.9% |
| 3 / 22 | DeepSeek V4 Pro 0813 · 1650.5B | 91.7% |
| 4 / 26 | DeepSeek V4 Flash 0731 · 304.2B | 91.0% |
| 5 / 27 | DeepSeek V4 Pro · 1598.8B | 90.9% |
| 6 / 28 | GLM 5.3 · 753.3B | 90.9% |
| 7 / 29 | MiniMax M3 · 427.0B | 90.9% |
| 8 / 31 | Kimi K2.6 · 1026.9B | 90.8% |
| 9 / 36 | GLM 5.3 Flash · 321.3B | 90.1% |
| 10 / 37 | GLM 5.1 · 753.9B | 89.9% |
| 11 / 47 | Inkling Small · 266.0B | 88.5% |
| 12 / 51 | Inkling · 952.4B | 88.3% |
| 13 / 55 | Kimi K2.7 Code · 1026.9B | 87.9% |
| 14 / 57 | GLM 5 · 753.9B | 87.8% |
| 15 / 59 | Kimi K2.5 · 1026.9B | 87.6% |
| 16 / 68 | Qwen3.5 397B A17B · 403.4B | 86.4% |
| 17 / 73 | Qwen3.6 27B · 27.8B | 85.9% |
| 18 / 83 | Qwen3.6 35B A3B · 36.0B | 84.9% |
| 19 / 84 | Kimi K2 Thinking · 1026.4B | 84.2% |
| 20 / 93 | GLM 4.7 · 358.3B | 83.3% |
| 21 / 112 | Qwen3 235B A22B Thinking 2507 · 235.1B | 80.0% |
| 22 / 115 | Qwen3.5 9B · 9.7B | 79.0% |
| 23 / 132 | DeepSeek R1 0528 · 684.5B | 76.3% |
| 24 / 137 | Gemma 4 31B IT · 31.3B | 75.8% |
| 25 / 138 | GPT OSS 120B · 116.8B | 75.8% |
| 26 / 151 | Gemma 4 26B A4B IT · 25.8B | 73.2% |
| 27 / 158 | Seed OSS 36B Instruct · 36.2B | 71.5% |
| 28 / 161 | Qwen3 235B A22B · 235.1B | 70.7% |
| 29 / 162 | Qwen3 30B A3B Thinking 2507 · 30.5B | 70.1% |
| 30 / 164 | DeepSeek R1 · 684.5B | 69.2% |
| 31 / 168 | DeepSeek v3 0324 · 684.5B | 67.6% |
| 32 / 171 | Llama 4 Maverick 17B 128E Instruct · 401.6B | 67.0% |
| 33 / 178 | Qwen3 32B · 32.8B | 65.7% |
| 34 / 181 | QwQ 32B · 32.8B | 65.3% |
| 35 / 182 | DeepSeek R1 Distill Qwen 32B · 32.8B | 64.1% |
| 36 / 185 | Qwen3 14B · 14.8B | 63.8% |
| 37 / 188 | Qwen3 30B A3B · 30.5B | 61.7% |
| 38 / 189 | GPT OSS 20B · 20.9B | 60.8% |
| 39 / 190 | GLM 4.7 Flash · 31.2B | 60.5% |
| 40 / 197 | Qwen3 8B · 8.2B | 56.8% |
| 41 / 198 | DeepSeek v3 · 684.5B | 56.5% |
| 42 / 200 | Magistral Small 2506 · 23.6B | 56.1% |
| 43 / 201 | Phi 4 · 14.7B | 56.1% |
| 44 / 202 | DeepSeek R1 Distill Llama 70B · 70.6B | 55.7% |
| 45 / 203 | Qwen3 30B A3B Instruct 2507 · 30.5B | 55.6% |
| 46 / 208 | Qwen3 4B · 4.0B | 52.3% |
| 47 / 209 | Llama 4 Scout 17B 16E Instruct · 108.6B | 51.8% |
| 48 / 211 | Llama 3.1 405B Instruct · 405.9B | 50.9% |
| 49 / 214 | Qwen2.5 72B Instruct · 72.7B | 49.1% |
| 50 / 222 | Gemma 3 27B IT · 27.4B | 47.7% |
| 51 / 225 | Llama 3.3 70B Instruct · 70.6B | 47.4% |
| 52 / 230 | Llama 3.1 Tulu 3 70B DPO · 70.6B | 46.3% |
| 53 / 231 | Qwen2.5 32B Instruct · 32.8B | 46.1% |
| 54 / 233 | Qwen3 4B Instruct 2507 · 4.0B | 45.8% |
| 55 / 235 | DeepSeek R1 Distill Qwen 14B · 14.8B | 44.7% |
| 56 / 236 | Llama 3.1 70B Instruct · 70.6B | 44.2% |
| 57 / 237 | WizardLM 2 8x22B · 140.6B | 43.4% |
| 58 / 242 | Llama 3.2 90B Vision Instruct · 88.6B | 41.0% |
| 59 / 243 | Qwen2 72B Instruct · 72.7B | 40.8% |
| 60 / 245 | Meta Llama 3 70B Instruct · 70.6B | 40.6% |
| 61 / 247 | Gemma 3 12B IT · 12.2B | 39.5% |
| 62 / 250 | Qwen3 1.7B · 2.0B | 38.0% |
| 63 / 252 | Hermes 2 Theta Llama 3 70B · 70.6B | 37.5% |
| 64 / 253 | Gemma 2 27B IT · 27.2B | 36.5% |
| 65 / 256 | Qwen2.5 7B Instruct · 7.6B | 35.5% |
| 66 / 260 | Eurus 2 7B PRIME · 7.6B | 33.9% |
| 67 / 261 | DeepSeek R1 Distill Qwen 1.5B · 1.8B | 33.6% |
| 68 / 265 | Yi 1.5 34B Chat · 34.4B | 32.0% |
| 69 / 266 | Qwen1.5 32B Chat · 32.5B | 30.7% |
| 70 / 268 | Mixtral 8x7B Instruct v0.1 · 46.7B | 30.6% |
| 71 / 271 | Qwen1.5 72B Chat · 72.3B | 28.8% |
| 72 / 272 | Granite 4.0 Micro · 3.4B | 28.3% |
| 73 / 275 | Gemma 2 9B IT · 9.2B | 27.5% |
| 74 / 278 | Llama 3.1 8B Instruct · 8.0B | 27.0% |
| 75 / 279 | Llama 2 70B Chat HF · 69.0B | 26.3% |
| 76 / 280 | Meta Llama 3 8B Instruct · 8.0B | 26.1% |
| 77 / 282 | Deepseek Llm 67B Chat · 67B | 24.6% |
| 78 / 284 | Llama 3.2 1B Instruct · 1.2B | 23.9% |
| 79 / 285 | Gemma 3 4B IT · 4.3B | 23.2% |
| 80 / 286 | Gemma 3 1B IT · 1000M | 20.0% |
| 81 / 287 | Mistral 7B Instruct v0.3 · 7.2B | 15.2% |
| 82 / 288 | Yi 34B Chat · 34.4B | 14.7% |
| 83 / 291 | DeepSeek R1 0528 Qwen3 8B · 8.2B | 9.3% |
Score vs model size
Which models give the most quality for their size — the ones worth running locally.
- Gemma 3 1B IT, 1000M, score 20.0% — on the efficiency frontier (best score at its size or smaller).
- Llama 3.2 1B Instruct, 1B, score 23.9% — on the efficiency frontier (best score at its size or smaller).
- DeepSeek R1 Distill Qwen 1.5B, 2B, score 33.6% — on the efficiency frontier (best score at its size or smaller).
- Qwen3 1.7B, 2B, score 38.0% — on the efficiency frontier (best score at its size or smaller).
- Qwen3 4B, 4B, score 52.3% — on the efficiency frontier (best score at its size or smaller).
- Qwen3 8B, 8B, score 56.8% — on the efficiency frontier (best score at its size or smaller).
- Qwen3.5 9B, 10B, score 79.0% — on the efficiency frontier (best score at its size or smaller).
- Qwen3.6 27B, 28B, score 85.9% — on the efficiency frontier (best score at its size or smaller).
- Inkling Small, 266B, score 88.5% — on the efficiency frontier (best score at its size or smaller).
- DeepSeek V4 Flash 0731, 304B, score 91.0% — on the efficiency frontier (best score at its size or smaller).
- GLM 5.2, 753B, score 91.9% — on the efficiency frontier (best score at its size or smaller).
- Kimi K3, 2.8T, score 93.1% — on the efficiency frontier (best score at its size or smaller).
GPQA Diamond: frequently asked questions
- What is the best open LLM on GPQA Diamond?
- Kimi K3 is the top open model on GPQA Diamond, scoring 93.1%. Among all models tested — including proprietary ones — it ranks #16. The top model overall is GPT 6 Astra Max (OpenAI) at 95.8%.
- What's the best GPQA Diamond model you can run on a 24 GB GPU?
- Qwen3.6 27B is the highest-scoring open model that fits in 24 GB at 4-bit quantization (about 15 GB), scoring 85.9% on GPQA Diamond.
- What's the best GPQA Diamond model you can run on a 12 GB GPU?
- Qwen3.5 9B is the highest-scoring open model that fits in 12 GB at 4-bit quantization (about 5 GB), scoring 79.0% on GPQA Diamond.
- Can open models match proprietary models on GPQA Diamond?
- Not quite on GPQA Diamond: the strongest proprietary model (GPT 6 Astra Max) scores 95.8%, ahead of the best open model (Kimi K3) at 93.1% — but you can run the open one yourself.
Scores aggregated from epoch. llmrun does not run this benchmark — see the source for methodology, or the about benchmarks for what it measures.