Reasoning
Chess Puzzles Leaderboard
Chess Puzzles is an Epoch AI evaluation that asks models to find the winning move in tactical chess positions drawn from real games. Because each puzzle has one verifiable correct answer and no natural-language shortcut, it isolates concrete multi-step planning.
Source: epoch65 open models ranked+140 proprietaryData through Sep 2026
All models ranked on Chess Puzzles
Proprietary / closed models are shown dimmed — you can't run them locally, but they show where the open field stands.
| # | Model | Score |
|---|---|---|
| 1 | GPT 6 Astra Max · proprietary | 72.0% |
| 2 | GPT 5.5 Pro Pre Release (xhigh) · proprietary | 64.0% |
| 3 | GPT 5.6 Sol Promax · proprietary | 64.0% |
| 4 | Gemini 3.8 Flash (high) · proprietary | 61.0% |
| 5 | GPT 5.4 Pro (Mar 05, 2026, xhigh) · proprietary | 58.6% |
| 6 | Gemini 3.1 Pro Preview · proprietary | 55.0% |
| 7 | GPT 5.6 Sol Max · proprietary | 55.0% |
| 8 | GPT 5.5 Pre Release (xhigh) · proprietary | 54.0% |
| 9 | GPT 5.6 Terra Max · proprietary | 54.0% |
| 10 | Gemini 3.5 Flash (high) · proprietary | 50.0% |
| 11 | Gemini 3.1 Pro Preview (high) · proprietary | 49.0% |
| 12 | GPT 5.2 (Dec 11, 2025, xhigh) · proprietary | 49.0% |
| 13 | Claude Fable 5.1 Max · proprietary | 47.0% |
| 14 | DeepSeek V4 Pro 0813 · 1650.5B | 47.0% |
| 15 | Gemini 3.7 Flash (high) · proprietary | 47.0% |
| 16 | Gemini 3.5 Flash (low) · proprietary | 45.0% |
| 17 | GPT 5.4 (Mar 05, 2026, xhigh) · proprietary | 44.0% |
| 18 | Gemini 3.5 Flash (minimal) · proprietary | 43.0% |
| 19 | Gemini 3.6 Flash (low) · proprietary | 43.0% |
| 20 | Claude Opus 5 Max · proprietary | 42.0% |
| 21 | Claude Fable 5 (high) · proprietary | 41.0% |
| 22 | Claude Fable 5 Max · proprietary | 41.0% |
| 23 | Gemini 3 Flash Preview (high) · proprietary | 40.0% |
| 24 | Gemini 3.6 Flash (high) · proprietary | 40.0% |
| 25 | GPT 5.2 (Dec 11, 2025, high) · proprietary | 40.0% |
| 26 | GPT 5.2 (Dec 11, 2025, medium) · proprietary | 40.0% |
| 27 | GPT 5.6 Luna Max · proprietary | 40.0% |
| 28 | Grok 4.6 (high) · proprietary | 40.0% |
| 29 | Qwen3.8 Max (Sep 02, xhigh) · proprietary | 40.0% |
| 30 | Kimi K3 · 2779.9B | 39.0% |
| 31 | Gemini 3 Flash Preview · proprietary | 38.0% |
| 32 | GPT 5.4 (Mar 05, 2026, high) · proprietary | 38.0% |
| 33 | GPT 5.4 (Mar 05, 2026, medium) · proprietary | 38.0% |
| 34 | Muse Spark 1.3 Max · proprietary | 38.0% |
| 35 | O3 (Apr 16, 2025, medium) · proprietary | 38.0% |
| 36 | GPT 5 (Aug 07, 2025, high) · proprietary | 37.0% |
| 37 | Grok 4.5 (high) · proprietary | 36.0% |
| 38 | Claude Sonnet 5 (xhigh) · proprietary | 35.0% |
| 39 | Gemini 3.6 Flash (minimal) · proprietary | 35.0% |
| 40 | Muse Spark 1.3 (xhigh) · proprietary | 35.0% |
| 41 | Claude Opus 4.8 Max · proprietary | 34.0% |
| 42 | O3 (Apr 16, 2025, high) · proprietary | 34.0% |
| 43 | Claude Opus 5 · proprietary | 33.0% |
| 44 | DeepSeek V4 Flash 0731 · 304.2B | 33.0% |
| 45 | GPT 5.1 (Nov 13, 2025, high) · proprietary | 32.0% |
| 46 | Gemini 3 Pro Preview · proprietary | 31.0% |
| 47 | Grok 4.6 (xhigh) · proprietary | 31.0% |
| 48 | Claude Opus 4.7 (xhigh) · proprietary | 30.0% |
| 49 | GPT 5 Mini (Aug 07, 2025, high) · proprietary | 30.0% |
| 50 | GPT 5.4 Nano (Mar 17, 2026, high) · proprietary | 30.0% |
| 51 | Claude Opus 4.8 (low) · proprietary | 29.0% |
| 52 | GPT 5 (Aug 07, 2025, medium) · proprietary | 29.0% |
| 53 | Qwen3.8 Max (xhigh) · proprietary | 29.0% |
| 54 | Claude Fable 5 (low) · proprietary | 28.0% |
| 55 | Grok 4 (Jul 09) · proprietary | 28.0% |
| 56 | GPT 5 Nano (Aug 07, 2025, high) · proprietary | 27.0% |
| 57 | GPT 5.6 Sol (low) · proprietary | 27.0% |
| 58 | O3 (Apr 16, 2025, low) · proprietary | 27.0% |
| 59 | GPT 5.5 (low) · proprietary | 26.0% |
| 60 | Kimi K2.6 · 1026.9B | 26.0% |
| 61 | O4 Mini (Apr 16, 2025, high) · proprietary | 26.0% |
| 62 | Qwen3.6 35B A3B · 36.0B | 26.0% |
| 63 | Gemini 3.1 Flash Lite (low) · proprietary | 25.0% |
| 64 | Grok 4.3 (high) · proprietary | 25.0% |
| 65 | Gemini 3.1 Flash Lite (minimal) · proprietary | 24.0% |
| 66 | GPT 5 (Aug 07, 2025, low) · proprietary | 24.0% |
| 67 | GPT 5.4 Mini (Mar 17, 2026, xhigh) · proprietary | 24.0% |
| 68 | Grok 4.20 0309 Reasoning · proprietary | 24.0% |
| 69 | Qwen3.7 Plus · proprietary | 24.0% |
| 70 | GPT 5.2 (Dec 11, 2025, low) · proprietary | 23.0% |
| 71 | Qwen3.7 Flash · proprietary | 23.0% |
| 72 | Gemini 3.5 Flash Lite (high) · proprietary | 22.0% |
| 73 | GPT 5.6 Terra (low) · proprietary | 22.0% |
| 74 | Qwen3.5 Plus · proprietary | 22.0% |
| 75 | Qwen3.6 27B · 27.8B | 22.0% |
| 76 | Gemini 3.5 Flash Lite (minimal) · proprietary | 21.0% |
| 77 | GLM 5.2 · 753.3B | 21.0% |
| 78 | GLM 5.3 · 753.3B | 21.0% |
| 79 | GPT 5.6 Luna (low) · proprietary | 21.0% |
| 80 | Inkling · 952.4B | 21.0% |
| 81 | Kimi K2.7 Code · 1026.9B | 21.0% |
| 82 | Qwen3.5 Flash · proprietary | 21.0% |
| 83 | Claude Opus 4.7 (low) · proprietary | 20.0% |
| 84 | Claude Opus 5 (low) · proprietary | 20.0% |
| 85 | DeepSeek V4 Pro · 1598.8B | 20.0% |
| 86 | Gemini 2.5 Pro · proprietary | 20.0% |
| 87 | Gemini 3.1 Flash Lite (high) · proprietary | 20.0% |
| 88 | GPT 5.4 (Mar 05, 2026, low) · proprietary | 20.0% |
| 89 | GPT OSS 120B · 116.8B | 20.0% |
| 90 | Kimi K2 Thinking · 1026.4B | 20.0% |
| 91 | O4 Mini (Apr 16, 2025, medium) · proprietary | 20.0% |
| 92 | Qwen3.6 Flash · proprietary | 20.0% |
| 93 | Qwen3.6 Max Preview · proprietary | 20.0% |
| 94 | GLM 5.1 · 753.9B | 19.0% |
| 95 | Qwen3.7 Max · proprietary | 19.0% |
| 96 | Gemini 3.5 Flash Lite (low) · proprietary | 18.0% |
| 97 | GPT 5.4 Mini (Mar 17, 2026, high) · proprietary | 18.0% |
| 98 | Inkling Small · 266.0B | 18.0% |
| 99 | Claude Opus 4.6 (32K) · proprietary | 17.0% |
| 100 | GPT 5.1 2025 11.13 None · proprietary | 17.0% |
| 101 | GPT 5.4 Nano (Mar 17, 2026, low) · proprietary | 17.0% |
| 102 | O3 Mini (Jan 31, 2025, high) · proprietary | 17.0% |
| 103 | Qwen3.6 Plus · proprietary | 17.0% |
| 104 | Claude Sonnet 5 Max · proprietary | 16.0% |
| 105 | GPT 5 (Aug 07, 2025, minimal) · proprietary | 16.0% |
| 106 | GPT 5 Nano (Aug 07, 2025, low) · proprietary | 15.0% |
| 107 | O1 (Dec 17, 2024, high) · proprietary | 15.0% |
| 108 | Claude Opus 4.6 Max · proprietary | 14.0% |
| 109 | DeepSeek Reasoner · proprietary | 14.0% |
| 110 | GLM 5.3 Flash · 321.3B | 14.0% |
| 111 | GPT 5.1 (Nov 13, 2025, low) · proprietary | 14.0% |
| 112 | MiniMax M3 · 427.0B | 14.0% |
| 113 | O4 Mini (Apr 16, 2025, low) · proprietary | 14.0% |
| 114 | Claude Opus 4.6 (120K) · proprietary | 13.0% |
| 115 | Claude Opus 4.8 None · proprietary | 13.0% |
| 116 | Claude Sonnet 4.6 (32K) · proprietary | 13.0% |
| 117 | GPT 4o (Aug 06, 2024) · proprietary | 13.0% |
| 118 | Qwen3.5 397B A17B · 403.4B | 13.0% |
| 119 | Seed OSS 36B Instruct · 36.2B | 13.0% |
| 120 | Claude Opus 4.5 (Nov 01, 2025, 32K) · proprietary | 12.0% |
| 121 | Claude Sonnet 4.5 (Sep 29, 2025, 32K) · proprietary | 12.0% |
| 122 | GPT 5 Mini (Aug 07, 2025, low) · proprietary | 12.0% |
| 123 | GPT 5.5 Instant · proprietary | 12.0% |
| 124 | Kimi K2.5 · 1026.9B | 12.0% |
| 125 | NVIDIA Nemotron 3 Ultra 550B A55B BF16 · 560.5B | 12.0% |
| 126 | O1 (Dec 17, 2024, medium) · proprietary | 12.0% |
| 127 | Qwen3 235B A22B Thinking 2507 · 235.1B | 12.0% |
| 128 | Qwen3.5 9B · 9.7B | 12.0% |
| 129 | Claude Opus 4.6 (64K) · proprietary | 10.0% |
| 130 | GLM 5 · 753.9B | 10.0% |
| 131 | GPT 5.5 None · proprietary | 10.0% |
| 132 | Qwen3.5 35B A3B · 36.0B | 10.0% |
| 133 | O3 Mini (Jan 31, 2025, medium) · proprietary | 9.0% |
| 134 | Qwen3.7 Flash None · proprietary | 9.0% |
| 135 | Qwen3.7 Plus None · proprietary | 9.0% |
| 136 | Claude Haiku 4.5 (Oct 01, 2025, 32K) · proprietary | 8.0% |
| 137 | Claude Sonnet 4.6 (medium) · proprietary | 8.0% |
| 138 | Qwen3 30B A3B Thinking 2507 · 30.5B | 8.0% |
| 139 | Claude Opus 4.1 (Aug 05, 2025) · proprietary | 7.0% |
| 140 | Claude Opus 4.7 Max · proprietary | 7.0% |
| 141 | GPT 4.1 Mini (Apr 14, 2025) · proprietary | 7.0% |
| 142 | GPT 5 Mini (Aug 07, 2025, minimal) · proprietary | 7.0% |
| 143 | GPT 5.6 Sol None · proprietary | 7.0% |
| 144 | O1 (Dec 17, 2024, low) · proprietary | 7.0% |
| 145 | Gemma 4 26B A4B IT · 25.8B | 6.0% |
| 146 | GLM 4.7 · 358.3B | 6.0% |
| 147 | GPT 4 Turbo (Apr 09, 2024) · proprietary | 6.0% |
| 148 | GPT 4.1 (Apr 14, 2025) · proprietary | 6.0% |
| 149 | O3 Mini (Jan 31, 2025, low) · proprietary | 6.0% |
| 150 | Claude 3 Opus (Feb 29, 2024) · proprietary | 5.0% |
| 151 | Claude Sonnet 4.6 (high) · proprietary | 5.0% |
| 152 | Gemma 4 31B IT · 31.3B | 5.0% |
| 153 | GPT 5.4 2026 03.05 None · proprietary | 5.0% |
| 154 | GPT 5.6 Terra None · proprietary | 5.0% |
| 155 | Qwen3 32B · 32.8B | 5.0% |
| 156 | Qwen3 8B · 8.2B | 5.0% |
| 157 | QwQ 32B · 32.8B | 5.0% |
| 158 | Claude Opus 4.5 (Nov 01, 2025) · proprietary | 4.0% |
| 159 | Claude Sonnet 4.5 (Sep 29, 2025) · proprietary | 4.0% |
| 160 | GPT 4 (Jun 13) · proprietary | 4.0% |
| 161 | GPT 5.2 2025 12.11 None · proprietary | 4.0% |
| 162 | GPT OSS 20B · 20.9B | 4.0% |
| 163 | Qwen3 14B · 14.8B | 4.0% |
| 164 | Qwen3 30B A3B · 30.5B | 4.0% |
| 165 | Qwen3 4B Instruct 2507 · 4.0B | 4.0% |
| 166 | Qwen3 Max (Sep 23, 2025) · proprietary | 4.0% |
| 167 | Claude Sonnet 4.6 Max · proprietary | 3.0% |
| 168 | DeepSeek R1 0528 Qwen3 8B · 8.2B | 3.0% |
| 169 | GPT 5.4 Mini 2026 03.17 None · proprietary | 3.0% |
| 170 | GPT 5.4 Nano 2026 03.17 None · proprietary | 3.0% |
| 171 | Magistral Small 2509 · 24.0B | 3.0% |
| 172 | GPT 5.6 Luna None · proprietary | 2.0% |
| 173 | Qwen3 30B A3B Instruct 2507 · 30.5B | 2.0% |
| 174 | DeepSeek Chat · proprietary | 1.0% |
| 175 | DeepSeek R1 Distill Qwen 14B · 14.8B | 1.0% |
| 176 | DeepSeek R1 Distill Qwen 32B · 32.8B | 1.0% |
| 177 | GPT 5 Nano (Aug 07, 2025, minimal) · proprietary | 1.0% |
| 178 | Mistral Small 3.1 24B Instruct 2503 · 24.0B | 1.0% |
| 179 | Mistral Small 3.2 24B Instruct 2506 · 24.0B | 1.0% |
| 180 | Phi 4 · 14.7B | 1.0% |
| 181 | Qwen3 4B · 4.0B | 1.0% |
| 182 | Deepseek Llm 67B Chat · 67B | 0.0% |
| 183 | DeepSeek R1 Distill Qwen 1.5B · 1.8B | 0.0% |
| 184 | Gemma 3 12B IT · 12.2B | 0.0% |
| 185 | Gemma 3 1B IT · 1000M | 0.0% |
| 186 | Gemma 3 27B IT · 27.4B | 0.0% |
| 187 | Gemma 3 4B IT · 4.3B | 0.0% |
| 188 | GLM 4.7 Flash · 31.2B | 0.0% |
| 189 | GPT 3.5 Turbo (Jan 25) · proprietary | 0.0% |
| 190 | GPT 4o Mini (Jul 18, 2024) · proprietary | 0.0% |
| 191 | Granite 4.0 1B · proprietary | 0.0% |
| 192 | Granite 4.0 350M · proprietary | 0.0% |
| 193 | Granite 4.0 Micro · 3.4B | 0.0% |
| 194 | Llama 2 13B Chat HF · 13.0B | 0.0% |
| 195 | Llama 2 7B Chat HF · 6.7B | 0.0% |
| 196 | Llama 3.1 8B Instruct · 8.0B | 0.0% |
| 197 | Llama 3.2 1B Instruct · 1.2B | 0.0% |
| 198 | Meta Llama 3 8B Instruct · 8.0B | 0.0% |
| 199 | Mistral 7B Instruct v0.3 · 7.2B | 0.0% |
| 200 | Mistral Small 24B Instruct 2501 · 23.6B | 0.0% |
| 201 | Phi 3 Mini 4k Instruct · 3.8B | 0.0% |
| 202 | Qwen2.5 32B Instruct · 32.8B | 0.0% |
| 203 | Qwen2.5 7B Instruct · 7.6B | 0.0% |
| 204 | Qwen3 1.7B · 2.0B | 0.0% |
| 205 | Qwen3.5 2B · 2.3B | 0.0% |
Score vs model size
Which models give the most quality for their size — the ones worth running locally.
- Gemma 3 1B IT, 1000M, score 0.0% — on the efficiency frontier (best score at its size or smaller).
- Qwen3 4B Instruct 2507, 4B, score 4.0% — on the efficiency frontier (best score at its size or smaller).
- Qwen3 8B, 8B, score 5.0% — on the efficiency frontier (best score at its size or smaller).
- Qwen3.5 9B, 10B, score 12.0% — on the efficiency frontier (best score at its size or smaller).
- Qwen3.6 27B, 28B, score 22.0% — on the efficiency frontier (best score at its size or smaller).
- Qwen3.6 35B A3B, 36B, score 26.0% — on the efficiency frontier (best score at its size or smaller).
- DeepSeek V4 Flash 0731, 304B, score 33.0% — on the efficiency frontier (best score at its size or smaller).
- DeepSeek V4 Pro 0813, 1.7T, score 47.0% — on the efficiency frontier (best score at its size or smaller).
Chess Puzzles: frequently asked questions
- What is the best open LLM on Chess Puzzles?
- DeepSeek V4 Pro 0813 is the top open model on Chess Puzzles, scoring 47.0%. Among all models tested — including proprietary ones — it ranks #13. The top model overall is GPT 6 Astra Max (OpenAI) at 72.0%.
- What's the best Chess Puzzles model you can run on a 24 GB GPU?
- Qwen3.6 35B A3B is the highest-scoring open model that fits in 24 GB at 4-bit quantization (about 20 GB), scoring 26.0% on Chess Puzzles.
- What's the best Chess Puzzles model you can run on a 12 GB GPU?
- Qwen3.5 9B is the highest-scoring open model that fits in 12 GB at 4-bit quantization (about 5 GB), scoring 12.0% on Chess Puzzles.
- Can open models match proprietary models on Chess Puzzles?
- Not quite on Chess Puzzles: the strongest proprietary model (GPT 6 Astra Max) scores 72.0%, ahead of the best open model (DeepSeek V4 Pro 0813) at 47.0% — but you can run the open one yourself.
Scores aggregated from epoch. llmrun does not run this benchmark — see the source for methodology, or the about benchmarks for what it measures.