Reasoning

Chess Puzzles Leaderboard

Chess Puzzles is an Epoch AI evaluation that asks models to find the winning move in tactical chess positions drawn from real games. Because each puzzle has one verifiable correct answer and no natural-language shortcut, it isolates concrete multi-step planning.

Source: epoch65 open models ranked+140 proprietaryData through Sep 2026

Open models ranked on Chess Puzzles

# shows rank among open models / rank overall (including proprietary).

#ModelScore
1 / 14DeepSeek V4 Pro 0813 · 1650.5B
47.0%
2 / 30Kimi K3 · 2779.9B
39.0%
3 / 44DeepSeek V4 Flash 0731 · 304.2B
33.0%
4 / 60Kimi K2.6 · 1026.9B
26.0%
5 / 62Qwen3.6 35B A3B · 36.0B
26.0%
6 / 75Qwen3.6 27B · 27.8B
22.0%
7 / 77GLM 5.2 · 753.3B
21.0%
8 / 78GLM 5.3 · 753.3B
21.0%
9 / 80Inkling · 952.4B
21.0%
10 / 81Kimi K2.7 Code · 1026.9B
21.0%
11 / 85DeepSeek V4 Pro · 1598.8B
20.0%
12 / 89GPT OSS 120B · 116.8B
20.0%
13 / 90Kimi K2 Thinking · 1026.4B
20.0%
14 / 94GLM 5.1 · 753.9B
19.0%
15 / 98Inkling Small · 266.0B
18.0%
16 / 110GLM 5.3 Flash · 321.3B
14.0%
17 / 112MiniMax M3 · 427.0B
14.0%
18 / 118Qwen3.5 397B A17B · 403.4B
13.0%
19 / 119Seed OSS 36B Instruct · 36.2B
13.0%
20 / 124Kimi K2.5 · 1026.9B
12.0%
21 / 125NVIDIA Nemotron 3 Ultra 550B A55B BF16 · 560.5B
12.0%
22 / 127Qwen3 235B A22B Thinking 2507 · 235.1B
12.0%
23 / 128Qwen3.5 9B · 9.7B
12.0%
24 / 130GLM 5 · 753.9B
10.0%
25 / 132Qwen3.5 35B A3B · 36.0B
10.0%
26 / 138Qwen3 30B A3B Thinking 2507 · 30.5B
8.0%
27 / 145Gemma 4 26B A4B IT · 25.8B
6.0%
28 / 146GLM 4.7 · 358.3B
6.0%
29 / 152Gemma 4 31B IT · 31.3B
5.0%
30 / 155Qwen3 32B · 32.8B
5.0%
31 / 156Qwen3 8B · 8.2B
5.0%
32 / 157QwQ 32B · 32.8B
5.0%
33 / 162GPT OSS 20B · 20.9B
4.0%
34 / 163Qwen3 14B · 14.8B
4.0%
35 / 164Qwen3 30B A3B · 30.5B
4.0%
36 / 165Qwen3 4B Instruct 2507 · 4.0B
4.0%
37 / 168DeepSeek R1 0528 Qwen3 8B · 8.2B
3.0%
38 / 171Magistral Small 2509 · 24.0B
3.0%
39 / 173Qwen3 30B A3B Instruct 2507 · 30.5B
2.0%
40 / 175DeepSeek R1 Distill Qwen 14B · 14.8B
1.0%
41 / 176DeepSeek R1 Distill Qwen 32B · 32.8B
1.0%
42 / 178Mistral Small 3.1 24B Instruct 2503 · 24.0B
1.0%
43 / 179Mistral Small 3.2 24B Instruct 2506 · 24.0B
1.0%
44 / 180Phi 4 · 14.7B
1.0%
45 / 181Qwen3 4B · 4.0B
1.0%
46 / 182Deepseek Llm 67B Chat · 67B
0.0%
47 / 183DeepSeek R1 Distill Qwen 1.5B · 1.8B
0.0%
48 / 184Gemma 3 12B IT · 12.2B
0.0%
49 / 185Gemma 3 1B IT · 1000M
0.0%
50 / 186Gemma 3 27B IT · 27.4B
0.0%
51 / 187Gemma 3 4B IT · 4.3B
0.0%
52 / 188GLM 4.7 Flash · 31.2B
0.0%
53 / 193Granite 4.0 Micro · 3.4B
0.0%
54 / 194Llama 2 13B Chat HF · 13.0B
0.0%
55 / 195Llama 2 7B Chat HF · 6.7B
0.0%
56 / 196Llama 3.1 8B Instruct · 8.0B
0.0%
57 / 197Llama 3.2 1B Instruct · 1.2B
0.0%
58 / 198Meta Llama 3 8B Instruct · 8.0B
0.0%
59 / 199Mistral 7B Instruct v0.3 · 7.2B
0.0%
60 / 200Mistral Small 24B Instruct 2501 · 23.6B
0.0%
61 / 201Phi 3 Mini 4k Instruct · 3.8B
0.0%
62 / 202Qwen2.5 32B Instruct · 32.8B
0.0%
63 / 203Qwen2.5 7B Instruct · 7.6B
0.0%
64 / 204Qwen3 1.7B · 2.0B
0.0%
65 / 205Qwen3.5 2B · 2.3B
0.0%

Score vs model size

Which models give the most quality for their size — the ones worth running locally.

1B10B100B1Tmodel size (log scale) →47.0%0.0%Kimi K3 · 2.8T · 39.0%Kimi K2.6 · 1T · 26.0%Kimi K2.7 Code · 1T · 21.0%Inkling · 952B · 21.0%GLM 5.2 · 753B · 21.0%GLM 5.3 · 753B · 21.0%DeepSeek V4 Pro · 1.6T · 20.0%Kimi K2 Thinking · 1T · 20.0%GPT OSS 120B · 117B · 20.0%GLM 5.1 · 754B · 19.0%Inkling Small · 266B · 18.0%MiniMax M3 · 427B · 14.0%GLM 5.3 Flash · 321B · 14.0%Seed OSS 36B Instruct · 36B · 13.0%Qwen3.5 397B A17B · 403B · 13.0%Kimi K2.5 · 1T · 12.0%NVIDIA Nemotron 3 Ultra 550B A55B BF16 · 561B · 12.0%Qwen3 235B A22B Thinking 2507 · 235B · 12.0%Qwen3.5 35B A3B · 36B · 10.0%GLM 5 · 754B · 10.0%Qwen3 30B A3B Thinking 2507 · 31B · 8.0%Gemma 4 26B A4B IT · 26B · 6.0%GLM 4.7 · 358B · 6.0%Gemma 4 31B IT · 31B · 5.0%Qwen3 32B · 33B · 5.0%QwQ 32B · 33B · 5.0%GPT OSS 20B · 21B · 4.0%Qwen3 14B · 15B · 4.0%Qwen3 30B A3B · 31B · 4.0%DeepSeek R1 0528 Qwen3 8B · 8B · 3.0%Magistral Small 2509 · 24B · 3.0%Qwen3 30B A3B Instruct 2507 · 31B · 2.0%DeepSeek R1 Distill Qwen 14B · 15B · 1.0%DeepSeek R1 Distill Qwen 32B · 33B · 1.0%Phi 4 · 15B · 1.0%Mistral Small 3.1 24B Instruct 2503 · 24B · 1.0%Mistral Small 3.2 24B Instruct 2506 · 24B · 1.0%Qwen3 4B · 4B · 1.0%Deepseek Llm 67B Chat · 67B · 0.0%DeepSeek R1 Distill Qwen 1.5B · 2B · 0.0%Gemma 3 12B IT · 12B · 0.0%Gemma 3 27B IT · 27B · 0.0%Gemma 3 4B IT · 4B · 0.0%Granite 4.0 Micro · 3B · 0.0%Llama 2 13B Chat HF · 13B · 0.0%Llama 2 7B Chat HF · 7B · 0.0%Llama 3.1 8B Instruct · 8B · 0.0%Llama 3.2 1B Instruct · 1B · 0.0%Meta Llama 3 8B Instruct · 8B · 0.0%Phi 3 Mini 4k Instruct · 4B · 0.0%Mistral 7B Instruct v0.3 · 7B · 0.0%Mistral Small 24B Instruct 2501 · 24B · 0.0%Qwen2.5 32B Instruct · 33B · 0.0%Qwen2.5 7B Instruct · 8B · 0.0%Qwen3 1.7B · 2B · 0.0%Qwen3.5 2B · 2B · 0.0%GLM 4.7 Flash · 31B · 0.0%Gemma 3 1B IT · 1000M · 0.0%Gemma 3 1B ITQwen3 4B Instruct 2507 · 4B · 4.0%Qwen3 4B Instruct 2507Qwen3 8B · 8B · 5.0%Qwen3 8BQwen3.5 9B · 10B · 12.0%Qwen3.5 9BQwen3.6 27B · 28B · 22.0%Qwen3.6 27BQwen3.6 35B A3B · 36B · 26.0%Qwen3.6 35B A3BDeepSeek V4 Flash 0731 · 304B · 33.0%DeepSeek V4 Flash 0731DeepSeek V4 Pro 0813 · 1.7T · 47.0%DeepSeek V4 Pro 0813
Each dot is a model. Up = higher score, left = smaller (easier to run locally). The dashed line marks the efficiency frontier — the best score you can get at each size or smaller.
  • Gemma 3 1B IT, 1000M, score 0.0% — on the efficiency frontier (best score at its size or smaller).
  • Qwen3 4B Instruct 2507, 4B, score 4.0% — on the efficiency frontier (best score at its size or smaller).
  • Qwen3 8B, 8B, score 5.0% — on the efficiency frontier (best score at its size or smaller).
  • Qwen3.5 9B, 10B, score 12.0% — on the efficiency frontier (best score at its size or smaller).
  • Qwen3.6 27B, 28B, score 22.0% — on the efficiency frontier (best score at its size or smaller).
  • Qwen3.6 35B A3B, 36B, score 26.0% — on the efficiency frontier (best score at its size or smaller).
  • DeepSeek V4 Flash 0731, 304B, score 33.0% — on the efficiency frontier (best score at its size or smaller).
  • DeepSeek V4 Pro 0813, 1.7T, score 47.0% — on the efficiency frontier (best score at its size or smaller).

Chess Puzzles: frequently asked questions

What is the best open LLM on Chess Puzzles?
DeepSeek V4 Pro 0813 is the top open model on Chess Puzzles, scoring 47.0%. Among all models tested — including proprietary ones — it ranks #13. The top model overall is GPT 6 Astra Max (OpenAI) at 72.0%.
What's the best Chess Puzzles model you can run on a 24 GB GPU?
Qwen3.6 35B A3B is the highest-scoring open model that fits in 24 GB at 4-bit quantization (about 20 GB), scoring 26.0% on Chess Puzzles.
What's the best Chess Puzzles model you can run on a 12 GB GPU?
Qwen3.5 9B is the highest-scoring open model that fits in 12 GB at 4-bit quantization (about 5 GB), scoring 12.0% on Chess Puzzles.
Can open models match proprietary models on Chess Puzzles?
Not quite on Chess Puzzles: the strongest proprietary model (GPT 6 Astra Max) scores 72.0%, ahead of the best open model (DeepSeek V4 Pro 0813) at 47.0% — but you can run the open one yourself.

Scores aggregated from epoch. llmrun does not run this benchmark — see the source for methodology, or the about benchmarks for what it measures.