Reasoning

Chess Puzzles Leaderboard

Chess Puzzles is an Epoch AI evaluation that asks models to find the winning move in tactical chess positions drawn from real games. Because each puzzle has one verifiable correct answer and no natural-language shortcut, it isolates concrete multi-step planning.

Source: epoch65 open models ranked+140 proprietaryData through Sep 2026

All models ranked on Chess Puzzles

Proprietary / closed models are shown dimmed — you can't run them locally, but they show where the open field stands.

#ModelScore
1GPT 6 Astra Max · proprietary
72.0%
2GPT 5.5 Pro Pre Release (xhigh) · proprietary
64.0%
3GPT 5.6 Sol Promax · proprietary
64.0%
4Gemini 3.8 Flash (high) · proprietary
61.0%
5GPT 5.4 Pro (Mar 05, 2026, xhigh) · proprietary
58.6%
6Gemini 3.1 Pro Preview · proprietary
55.0%
7GPT 5.6 Sol Max · proprietary
55.0%
8GPT 5.5 Pre Release (xhigh) · proprietary
54.0%
9GPT 5.6 Terra Max · proprietary
54.0%
10Gemini 3.5 Flash (high) · proprietary
50.0%
11Gemini 3.1 Pro Preview (high) · proprietary
49.0%
12GPT 5.2 (Dec 11, 2025, xhigh) · proprietary
49.0%
13Claude Fable 5.1 Max · proprietary
47.0%
14DeepSeek V4 Pro 0813 · 1650.5B
47.0%
15Gemini 3.7 Flash (high) · proprietary
47.0%
16Gemini 3.5 Flash (low) · proprietary
45.0%
17GPT 5.4 (Mar 05, 2026, xhigh) · proprietary
44.0%
18Gemini 3.5 Flash (minimal) · proprietary
43.0%
19Gemini 3.6 Flash (low) · proprietary
43.0%
20Claude Opus 5 Max · proprietary
42.0%
21Claude Fable 5 (high) · proprietary
41.0%
22Claude Fable 5 Max · proprietary
41.0%
23Gemini 3 Flash Preview (high) · proprietary
40.0%
24Gemini 3.6 Flash (high) · proprietary
40.0%
25GPT 5.2 (Dec 11, 2025, high) · proprietary
40.0%
26GPT 5.2 (Dec 11, 2025, medium) · proprietary
40.0%
27GPT 5.6 Luna Max · proprietary
40.0%
28Grok 4.6 (high) · proprietary
40.0%
29Qwen3.8 Max (Sep 02, xhigh) · proprietary
40.0%
30Kimi K3 · 2779.9B
39.0%
31Gemini 3 Flash Preview · proprietary
38.0%
32GPT 5.4 (Mar 05, 2026, high) · proprietary
38.0%
33GPT 5.4 (Mar 05, 2026, medium) · proprietary
38.0%
34Muse Spark 1.3 Max · proprietary
38.0%
35O3 (Apr 16, 2025, medium) · proprietary
38.0%
36GPT 5 (Aug 07, 2025, high) · proprietary
37.0%
37Grok 4.5 (high) · proprietary
36.0%
38Claude Sonnet 5 (xhigh) · proprietary
35.0%
39Gemini 3.6 Flash (minimal) · proprietary
35.0%
40Muse Spark 1.3 (xhigh) · proprietary
35.0%
41Claude Opus 4.8 Max · proprietary
34.0%
42O3 (Apr 16, 2025, high) · proprietary
34.0%
43Claude Opus 5 · proprietary
33.0%
44DeepSeek V4 Flash 0731 · 304.2B
33.0%
45GPT 5.1 (Nov 13, 2025, high) · proprietary
32.0%
46Gemini 3 Pro Preview · proprietary
31.0%
47Grok 4.6 (xhigh) · proprietary
31.0%
48Claude Opus 4.7 (xhigh) · proprietary
30.0%
49GPT 5 Mini (Aug 07, 2025, high) · proprietary
30.0%
50GPT 5.4 Nano (Mar 17, 2026, high) · proprietary
30.0%
51Claude Opus 4.8 (low) · proprietary
29.0%
52GPT 5 (Aug 07, 2025, medium) · proprietary
29.0%
53Qwen3.8 Max (xhigh) · proprietary
29.0%
54Claude Fable 5 (low) · proprietary
28.0%
55Grok 4 (Jul 09) · proprietary
28.0%
56GPT 5 Nano (Aug 07, 2025, high) · proprietary
27.0%
57GPT 5.6 Sol (low) · proprietary
27.0%
58O3 (Apr 16, 2025, low) · proprietary
27.0%
59GPT 5.5 (low) · proprietary
26.0%
60Kimi K2.6 · 1026.9B
26.0%
61O4 Mini (Apr 16, 2025, high) · proprietary
26.0%
62Qwen3.6 35B A3B · 36.0B
26.0%
63Gemini 3.1 Flash Lite (low) · proprietary
25.0%
64Grok 4.3 (high) · proprietary
25.0%
65Gemini 3.1 Flash Lite (minimal) · proprietary
24.0%
66GPT 5 (Aug 07, 2025, low) · proprietary
24.0%
67GPT 5.4 Mini (Mar 17, 2026, xhigh) · proprietary
24.0%
68Grok 4.20 0309 Reasoning · proprietary
24.0%
69Qwen3.7 Plus · proprietary
24.0%
70GPT 5.2 (Dec 11, 2025, low) · proprietary
23.0%
71Qwen3.7 Flash · proprietary
23.0%
72Gemini 3.5 Flash Lite (high) · proprietary
22.0%
73GPT 5.6 Terra (low) · proprietary
22.0%
74Qwen3.5 Plus · proprietary
22.0%
75Qwen3.6 27B · 27.8B
22.0%
76Gemini 3.5 Flash Lite (minimal) · proprietary
21.0%
77GLM 5.2 · 753.3B
21.0%
78GLM 5.3 · 753.3B
21.0%
79GPT 5.6 Luna (low) · proprietary
21.0%
80Inkling · 952.4B
21.0%
81Kimi K2.7 Code · 1026.9B
21.0%
82Qwen3.5 Flash · proprietary
21.0%
83Claude Opus 4.7 (low) · proprietary
20.0%
84Claude Opus 5 (low) · proprietary
20.0%
85DeepSeek V4 Pro · 1598.8B
20.0%
86Gemini 2.5 Pro · proprietary
20.0%
87Gemini 3.1 Flash Lite (high) · proprietary
20.0%
88GPT 5.4 (Mar 05, 2026, low) · proprietary
20.0%
89GPT OSS 120B · 116.8B
20.0%
90Kimi K2 Thinking · 1026.4B
20.0%
91O4 Mini (Apr 16, 2025, medium) · proprietary
20.0%
92Qwen3.6 Flash · proprietary
20.0%
93Qwen3.6 Max Preview · proprietary
20.0%
94GLM 5.1 · 753.9B
19.0%
95Qwen3.7 Max · proprietary
19.0%
96Gemini 3.5 Flash Lite (low) · proprietary
18.0%
97GPT 5.4 Mini (Mar 17, 2026, high) · proprietary
18.0%
98Inkling Small · 266.0B
18.0%
99Claude Opus 4.6 (32K) · proprietary
17.0%
100GPT 5.1 2025 11.13 None · proprietary
17.0%
101GPT 5.4 Nano (Mar 17, 2026, low) · proprietary
17.0%
102O3 Mini (Jan 31, 2025, high) · proprietary
17.0%
103Qwen3.6 Plus · proprietary
17.0%
104Claude Sonnet 5 Max · proprietary
16.0%
105GPT 5 (Aug 07, 2025, minimal) · proprietary
16.0%
106GPT 5 Nano (Aug 07, 2025, low) · proprietary
15.0%
107O1 (Dec 17, 2024, high) · proprietary
15.0%
108Claude Opus 4.6 Max · proprietary
14.0%
109DeepSeek Reasoner · proprietary
14.0%
110GLM 5.3 Flash · 321.3B
14.0%
111GPT 5.1 (Nov 13, 2025, low) · proprietary
14.0%
112MiniMax M3 · 427.0B
14.0%
113O4 Mini (Apr 16, 2025, low) · proprietary
14.0%
114Claude Opus 4.6 (120K) · proprietary
13.0%
115Claude Opus 4.8 None · proprietary
13.0%
116Claude Sonnet 4.6 (32K) · proprietary
13.0%
117GPT 4o (Aug 06, 2024) · proprietary
13.0%
118Qwen3.5 397B A17B · 403.4B
13.0%
119Seed OSS 36B Instruct · 36.2B
13.0%
120Claude Opus 4.5 (Nov 01, 2025, 32K) · proprietary
12.0%
121Claude Sonnet 4.5 (Sep 29, 2025, 32K) · proprietary
12.0%
122GPT 5 Mini (Aug 07, 2025, low) · proprietary
12.0%
123GPT 5.5 Instant · proprietary
12.0%
124Kimi K2.5 · 1026.9B
12.0%
125NVIDIA Nemotron 3 Ultra 550B A55B BF16 · 560.5B
12.0%
126O1 (Dec 17, 2024, medium) · proprietary
12.0%
127Qwen3 235B A22B Thinking 2507 · 235.1B
12.0%
128Qwen3.5 9B · 9.7B
12.0%
129Claude Opus 4.6 (64K) · proprietary
10.0%
130GLM 5 · 753.9B
10.0%
131GPT 5.5 None · proprietary
10.0%
132Qwen3.5 35B A3B · 36.0B
10.0%
133O3 Mini (Jan 31, 2025, medium) · proprietary
9.0%
134Qwen3.7 Flash None · proprietary
9.0%
135Qwen3.7 Plus None · proprietary
9.0%
136Claude Haiku 4.5 (Oct 01, 2025, 32K) · proprietary
8.0%
137Claude Sonnet 4.6 (medium) · proprietary
8.0%
138Qwen3 30B A3B Thinking 2507 · 30.5B
8.0%
139Claude Opus 4.1 (Aug 05, 2025) · proprietary
7.0%
140Claude Opus 4.7 Max · proprietary
7.0%
141GPT 4.1 Mini (Apr 14, 2025) · proprietary
7.0%
142GPT 5 Mini (Aug 07, 2025, minimal) · proprietary
7.0%
143GPT 5.6 Sol None · proprietary
7.0%
144O1 (Dec 17, 2024, low) · proprietary
7.0%
145Gemma 4 26B A4B IT · 25.8B
6.0%
146GLM 4.7 · 358.3B
6.0%
147GPT 4 Turbo (Apr 09, 2024) · proprietary
6.0%
148GPT 4.1 (Apr 14, 2025) · proprietary
6.0%
149O3 Mini (Jan 31, 2025, low) · proprietary
6.0%
150Claude 3 Opus (Feb 29, 2024) · proprietary
5.0%
151Claude Sonnet 4.6 (high) · proprietary
5.0%
152Gemma 4 31B IT · 31.3B
5.0%
153GPT 5.4 2026 03.05 None · proprietary
5.0%
154GPT 5.6 Terra None · proprietary
5.0%
155Qwen3 32B · 32.8B
5.0%
156Qwen3 8B · 8.2B
5.0%
157QwQ 32B · 32.8B
5.0%
158Claude Opus 4.5 (Nov 01, 2025) · proprietary
4.0%
159Claude Sonnet 4.5 (Sep 29, 2025) · proprietary
4.0%
160GPT 4 (Jun 13) · proprietary
4.0%
161GPT 5.2 2025 12.11 None · proprietary
4.0%
162GPT OSS 20B · 20.9B
4.0%
163Qwen3 14B · 14.8B
4.0%
164Qwen3 30B A3B · 30.5B
4.0%
165Qwen3 4B Instruct 2507 · 4.0B
4.0%
166Qwen3 Max (Sep 23, 2025) · proprietary
4.0%
167Claude Sonnet 4.6 Max · proprietary
3.0%
168DeepSeek R1 0528 Qwen3 8B · 8.2B
3.0%
169GPT 5.4 Mini 2026 03.17 None · proprietary
3.0%
170GPT 5.4 Nano 2026 03.17 None · proprietary
3.0%
171Magistral Small 2509 · 24.0B
3.0%
172GPT 5.6 Luna None · proprietary
2.0%
173Qwen3 30B A3B Instruct 2507 · 30.5B
2.0%
174DeepSeek Chat · proprietary
1.0%
175DeepSeek R1 Distill Qwen 14B · 14.8B
1.0%
176DeepSeek R1 Distill Qwen 32B · 32.8B
1.0%
177GPT 5 Nano (Aug 07, 2025, minimal) · proprietary
1.0%
178Mistral Small 3.1 24B Instruct 2503 · 24.0B
1.0%
179Mistral Small 3.2 24B Instruct 2506 · 24.0B
1.0%
180Phi 4 · 14.7B
1.0%
181Qwen3 4B · 4.0B
1.0%
182Deepseek Llm 67B Chat · 67B
0.0%
183DeepSeek R1 Distill Qwen 1.5B · 1.8B
0.0%
184Gemma 3 12B IT · 12.2B
0.0%
185Gemma 3 1B IT · 1000M
0.0%
186Gemma 3 27B IT · 27.4B
0.0%
187Gemma 3 4B IT · 4.3B
0.0%
188GLM 4.7 Flash · 31.2B
0.0%
189GPT 3.5 Turbo (Jan 25) · proprietary
0.0%
190GPT 4o Mini (Jul 18, 2024) · proprietary
0.0%
191Granite 4.0 1B · proprietary
0.0%
192Granite 4.0 350M · proprietary
0.0%
193Granite 4.0 Micro · 3.4B
0.0%
194Llama 2 13B Chat HF · 13.0B
0.0%
195Llama 2 7B Chat HF · 6.7B
0.0%
196Llama 3.1 8B Instruct · 8.0B
0.0%
197Llama 3.2 1B Instruct · 1.2B
0.0%
198Meta Llama 3 8B Instruct · 8.0B
0.0%
199Mistral 7B Instruct v0.3 · 7.2B
0.0%
200Mistral Small 24B Instruct 2501 · 23.6B
0.0%
201Phi 3 Mini 4k Instruct · 3.8B
0.0%
202Qwen2.5 32B Instruct · 32.8B
0.0%
203Qwen2.5 7B Instruct · 7.6B
0.0%
204Qwen3 1.7B · 2.0B
0.0%
205Qwen3.5 2B · 2.3B
0.0%

Score vs model size

Which models give the most quality for their size — the ones worth running locally.

1B10B100B1Tmodel size (log scale) →47.0%0.0%Kimi K3 · 2.8T · 39.0%Kimi K2.6 · 1T · 26.0%Kimi K2.7 Code · 1T · 21.0%Inkling · 952B · 21.0%GLM 5.2 · 753B · 21.0%GLM 5.3 · 753B · 21.0%DeepSeek V4 Pro · 1.6T · 20.0%Kimi K2 Thinking · 1T · 20.0%GPT OSS 120B · 117B · 20.0%GLM 5.1 · 754B · 19.0%Inkling Small · 266B · 18.0%MiniMax M3 · 427B · 14.0%GLM 5.3 Flash · 321B · 14.0%Seed OSS 36B Instruct · 36B · 13.0%Qwen3.5 397B A17B · 403B · 13.0%Kimi K2.5 · 1T · 12.0%NVIDIA Nemotron 3 Ultra 550B A55B BF16 · 561B · 12.0%Qwen3 235B A22B Thinking 2507 · 235B · 12.0%Qwen3.5 35B A3B · 36B · 10.0%GLM 5 · 754B · 10.0%Qwen3 30B A3B Thinking 2507 · 31B · 8.0%Gemma 4 26B A4B IT · 26B · 6.0%GLM 4.7 · 358B · 6.0%Gemma 4 31B IT · 31B · 5.0%Qwen3 32B · 33B · 5.0%QwQ 32B · 33B · 5.0%GPT OSS 20B · 21B · 4.0%Qwen3 14B · 15B · 4.0%Qwen3 30B A3B · 31B · 4.0%DeepSeek R1 0528 Qwen3 8B · 8B · 3.0%Magistral Small 2509 · 24B · 3.0%Qwen3 30B A3B Instruct 2507 · 31B · 2.0%DeepSeek R1 Distill Qwen 14B · 15B · 1.0%DeepSeek R1 Distill Qwen 32B · 33B · 1.0%Phi 4 · 15B · 1.0%Mistral Small 3.1 24B Instruct 2503 · 24B · 1.0%Mistral Small 3.2 24B Instruct 2506 · 24B · 1.0%Qwen3 4B · 4B · 1.0%Deepseek Llm 67B Chat · 67B · 0.0%DeepSeek R1 Distill Qwen 1.5B · 2B · 0.0%Gemma 3 12B IT · 12B · 0.0%Gemma 3 27B IT · 27B · 0.0%Gemma 3 4B IT · 4B · 0.0%Granite 4.0 Micro · 3B · 0.0%Llama 2 13B Chat HF · 13B · 0.0%Llama 2 7B Chat HF · 7B · 0.0%Llama 3.1 8B Instruct · 8B · 0.0%Llama 3.2 1B Instruct · 1B · 0.0%Meta Llama 3 8B Instruct · 8B · 0.0%Phi 3 Mini 4k Instruct · 4B · 0.0%Mistral 7B Instruct v0.3 · 7B · 0.0%Mistral Small 24B Instruct 2501 · 24B · 0.0%Qwen2.5 32B Instruct · 33B · 0.0%Qwen2.5 7B Instruct · 8B · 0.0%Qwen3 1.7B · 2B · 0.0%Qwen3.5 2B · 2B · 0.0%GLM 4.7 Flash · 31B · 0.0%Gemma 3 1B IT · 1000M · 0.0%Gemma 3 1B ITQwen3 4B Instruct 2507 · 4B · 4.0%Qwen3 4B Instruct 2507Qwen3 8B · 8B · 5.0%Qwen3 8BQwen3.5 9B · 10B · 12.0%Qwen3.5 9BQwen3.6 27B · 28B · 22.0%Qwen3.6 27BQwen3.6 35B A3B · 36B · 26.0%Qwen3.6 35B A3BDeepSeek V4 Flash 0731 · 304B · 33.0%DeepSeek V4 Flash 0731DeepSeek V4 Pro 0813 · 1.7T · 47.0%DeepSeek V4 Pro 0813
Each dot is a model. Up = higher score, left = smaller (easier to run locally). The dashed line marks the efficiency frontier — the best score you can get at each size or smaller.
  • Gemma 3 1B IT, 1000M, score 0.0% — on the efficiency frontier (best score at its size or smaller).
  • Qwen3 4B Instruct 2507, 4B, score 4.0% — on the efficiency frontier (best score at its size or smaller).
  • Qwen3 8B, 8B, score 5.0% — on the efficiency frontier (best score at its size or smaller).
  • Qwen3.5 9B, 10B, score 12.0% — on the efficiency frontier (best score at its size or smaller).
  • Qwen3.6 27B, 28B, score 22.0% — on the efficiency frontier (best score at its size or smaller).
  • Qwen3.6 35B A3B, 36B, score 26.0% — on the efficiency frontier (best score at its size or smaller).
  • DeepSeek V4 Flash 0731, 304B, score 33.0% — on the efficiency frontier (best score at its size or smaller).
  • DeepSeek V4 Pro 0813, 1.7T, score 47.0% — on the efficiency frontier (best score at its size or smaller).

Chess Puzzles: frequently asked questions

What is the best open LLM on Chess Puzzles?
DeepSeek V4 Pro 0813 is the top open model on Chess Puzzles, scoring 47.0%. Among all models tested — including proprietary ones — it ranks #13. The top model overall is GPT 6 Astra Max (OpenAI) at 72.0%.
What's the best Chess Puzzles model you can run on a 24 GB GPU?
Qwen3.6 35B A3B is the highest-scoring open model that fits in 24 GB at 4-bit quantization (about 20 GB), scoring 26.0% on Chess Puzzles.
What's the best Chess Puzzles model you can run on a 12 GB GPU?
Qwen3.5 9B is the highest-scoring open model that fits in 12 GB at 4-bit quantization (about 5 GB), scoring 12.0% on Chess Puzzles.
Can open models match proprietary models on Chess Puzzles?
Not quite on Chess Puzzles: the strongest proprietary model (GPT 6 Astra Max) scores 72.0%, ahead of the best open model (DeepSeek V4 Pro 0813) at 47.0% — but you can run the open one yourself.

Scores aggregated from epoch. llmrun does not run this benchmark — see the source for methodology, or the about benchmarks for what it measures.