Reasoning

DTBench Leaderboard

DTBench is a set of handcrafted, expert-validated multiple-choice questions on decision theory — Newcomb-like problems and related puzzles — built by the Conceptual Reasoning Index project with Anthropic. It tests whether a model can follow an abstract argument rather than recall a fact.

Source: epoch67 open models ranked+143 proprietaryData through Sep 2026

All models ranked on DTBench

Proprietary / closed models are shown dimmed — you can't run them locally, but they show where the open field stands.

#ModelScore
1Claude Fable 5 Max · proprietary
98.4%
2Claude Opus 5 (high) · proprietary
97.9%
3Claude Opus 5 (xhigh) · proprietary
97.9%
4Claude Fable 5.1 (unspecified) · proprietary
97.6%
5Claude Fable 5.1 (xhigh) · proprietary
97.6%
6Claude Opus 5 Max · proprietary
97.6%
7Claude Fable 5 (xhigh) · proprietary
97.3%
8GPT 6 Astra (xhigh) · proprietary
97.3%
9Grok 4.6 (xhigh) · proprietary
97.3%
10Claude Fable 5.1 (medium) · proprietary
97.1%
11Gemini 3.1 Pro Preview (medium) · proprietary
97.1%
12GPT 6 Astra (medium) · proprietary
97.1%
13GPT 6 Astra (unspecified) · proprietary
97.1%
14Gemini 3.7 Flash (high) · proprietary
96.8%
15Claude Fable 5 (high) · proprietary
96.5%
16Claude Fable 5.1 (high) · proprietary
96.5%
17GPT 6 Astra (high) · proprietary
96.5%
18Grok 4.5 (high) · proprietary
96.5%
19Muse Spark 1.3 (unspecified) · proprietary
96.5%
20GPT 5.5 (xhigh) · proprietary
96.0%
21GPT 5.5 Pro (xhigh) · proprietary
96.0%
22GPT 5.6 Sol Promax · proprietary
96.0%
23Claude Fable 5 (medium) · proprietary
95.7%
24GPT 5.6 Sol (xhigh) · proprietary
95.7%
25Gemini 3.8 Flash (unspecified) · proprietary
95.7%
26Gemini 3.1 Pro Preview (high) · proprietary
95.5%
27Gemini 3.6 Flash (high) · proprietary
95.5%
28GPT 5.6 Sol (high) · proprietary
95.5%
29GPT 5.6 Sol Max · proprietary
95.5%
30GPT 5.6 Sol (medium) · proprietary
95.2%
31Claude Opus 4.8 Max · proprietary
94.9%
32Claude Opus 5 (medium) · proprietary
94.9%
33Gemini 3.1 Pro Preview (low) · proprietary
94.9%
34Gemini 3.8 Flash (low) · proprietary
94.9%
35Gemini 3.8 Flash (medium) · proprietary
94.9%
36GPT 6 Astra (low) · proprietary
94.9%
37Gemini 3.5 Flash (high) · proprietary
94.7%
38Gemini 3.7 Flash (medium) · proprietary
94.7%
39Muse Spark 1.2 (xhigh) · proprietary
94.7%
40Claude Opus 4.7 Max · proprietary
94.7%
41GPT 5.4 (Mar 05, 2026, xhigh) · proprietary
94.4%
42Muse Spark 1.1 (high) · proprietary
94.4%
43Claude Fable 5.1 (low) · proprietary
94.1%
44DeepSeek V4 Pro 0813 · 1650.5B
93.9%
45GPT 5.6 Sol (low) · proprietary
93.9%
46GLM 5.2 · 753.3B
93.6%
47Gemini 3.7 Flash (low) · proprietary
93.3%
48GPT 5.6 Terra Max · proprietary
93.3%
49Claude Opus 5 (low) · proprietary
93.1%
50GPT 5.6 Terra (xhigh) · proprietary
92.5%
51Claude Sonnet 5 Max · proprietary
92.5%
52Qwen3.7 Max · proprietary
92.3%
53Qwen3.8 Max (xhigh) · proprietary
92.0%
54Claude Fable 5 (low) · proprietary
91.5%
55Claude Opus 4.6 Max · proprietary
91.2%
56GPT 5.6 Terra (high) · proprietary
91.2%
57Kimi K3 · 2779.9B
91.2%
58DeepSeek V4 Flash 0731 · 304.2B
90.9%
59GPT 5.2 (Dec 11, 2025, xhigh) · proprietary
90.9%
60Kimi K2.6 · 1026.9B
90.9%
61DeepSeek V4 Pro · 1598.8B
90.7%
62GPT 5 (Aug 07, 2025, high) · proprietary
90.7%
63Grok 4.3 (high) · proprietary
90.7%
64GPT 5.1 (Nov 13, 2025, high) · proprietary
90.1%
65GPT 5.6 Terra (low) · proprietary
90.1%
66GPT 5.6 Terra (medium) · proprietary
90.1%
67Grok 4.20 · proprietary
90.1%
68NVIDIA Nemotron 3 Ultra 550B A55B BF16 · 560.5B
90.1%
69Claude Opus 4.5 (Nov 01, 2025, high) · proprietary
89.9%
70Claude Sonnet 4.6 Max · proprietary
89.9%
71DeepSeek V4.1 Flash · 763.2B
89.9%
72Claude Sonnet 5 (xhigh) · proprietary
89.5%
73Gemini 3 Flash Preview (high) · proprietary
89.1%
74GPT 5.6 Luna (xhigh) · proprietary
89.1%
75GPT 5.6 Luna Max · proprietary
88.8%
76Qwen3.8 27B · 27.8B
88.0%
77DeepSeek V3.2 Exp · 685.4B
87.7%
78Grok 4.1 Fast Reasoning · proprietary
87.7%
79GLM 5.3 · 753.3B
87.7%
80Inkling · 952.4B
87.5%
81Qwen3.5 397B A17B · 403.4B
87.5%
82Qwen3.6 Max Preview · proprietary
87.2%
83O3 Pro (Jun 10, 2025, high) · proprietary
86.9%
84GPT 5.6 Luna (high) · proprietary
86.7%
85DeepSeek V4 Flash · 290.9B
86.4%
86DeepSeek V3.2 · 685.4B
85.6%
87O3 (Apr 16, 2025, high) · proprietary
84.8%
88Claude Sonnet 5 (high) · proprietary
84.5%
89MiMo V2.5 Pro · 1023.2B
84.5%
90Qwen3.5 122B A10B · 125.1B
84.3%
91Qwen3.7 Plus · proprietary
84.0%
92GPT 5.6 Sol None · proprietary
83.7%
93Gemini 3.5 Flash Lite (high) · proprietary
83.5%
94Claude Sonnet 4.5 (Sep 29, 2025, unspecified) · proprietary
83.2%
95GPT 5.6 Luna (medium) · proprietary
83.2%
96Qwen3.5 Flash · proprietary
82.9%
97DeepSeek V3.1 · 684.5B
82.7%
98Gemma 4 31B IT · 31.3B
82.7%
99Grok 4 Fast · proprietary
82.7%
100Gemini 2.5 Pro · proprietary
82.4%
101Qwen3.5 27B · 27.8B
82.4%
102Qwen3 Max (Sep 23, 2025) · proprietary
82.1%
103Qwen3.6 Plus · proprietary
81.9%
104Claude Opus 4 (May 14, 2025, unspecified) · proprietary
81.6%
105DeepSeek V3.1 Terminus · 684.5B
81.3%
106Qwen Plus (Apr 28, 2025) · proprietary
81.1%
107GPT 5 Mini (Aug 07, 2025, high) · proprietary
80.5%
108Qwen3.5 Plus · proprietary
80.5%
109GPT 5.4 Nano (Mar 17, 2026, xhigh) · proprietary
80.3%
110Qwen3 235B A22B Thinking 2507 · 235.1B
80.3%
111Claude Opus 4.1 (Aug 05, 2025, unspecified) · proprietary
80.0%
112GPT 5.4 Mini (Mar 17, 2026, xhigh) · proprietary
80.0%
113GPT 5.6 Luna (low) · proprietary
80.0%
114Qwen3.5 35B A3B · 36.0B
80.0%
115Claude Sonnet 5 (medium) · proprietary
79.2%
116MiniMax M3 · 427.0B
78.9%
117GPT 5.6 Terra None · proprietary
78.7%
118Qwen3 235B A22B Instruct 2507 · 235.1B
78.4%
119Qwen3.6 27B · 27.8B
78.1%
120O4 Mini (Apr 16, 2025, high) · proprietary
77.6%
121Claude Sonnet 4 (May 14, 2025, unspecified) · proprietary
77.1%
122Qwen3.6 Flash · proprietary
77.1%
123Gemini 3.1 Flash Lite (high) · proprietary
76.8%
124Gemini 2.5 Flash · proprietary
76.5%
125GPT OSS 120B · 116.8B
76.3%
126Qwen3 235B A22B · 235.1B
75.7%
127Mistral Medium 2604 · proprietary
75.5%
128Gemma 4 26B A4B IT · 25.8B
74.9%
129O1 (Dec 17, 2024, high) · proprietary
74.7%
130Qwen3.6 35B A3B · 36.0B
73.9%
131Claude Haiku 4.5 (Oct 01, 2025, unspecified) · proprietary
73.6%
132Claude Sonnet 5 (low) · proprietary
73.6%
133Qwen3.5 9B · 9.7B
71.2%
134Mistral Small 2603 · proprietary
70.9%
135Qwen3 30B A3B Thinking 2507 · 30.5B
69.3%
136GPT 5.6 Luna None · proprietary
69.3%
137GPT 4.1 Mini (Apr 14, 2025) · proprietary
68.8%
138O3 Mini (Jan 31, 2025, high) · proprietary
68.8%
139GPT 4.1 (Apr 14, 2025) · proprietary
68.3%
140GPT OSS 20B · 20.9B
68.0%
141Claude 3.5 Sonnet (Oct 22, 2024) · proprietary
67.8%
142Claude 3.5 Sonnet (Jun 20, 2024) · proprietary
67.8%
143Mistral Medium 2508 · proprietary
67.5%
144Qwen3 32B · 32.8B
67.5%
145Qwen3 30B A3B Instruct 2507 · 30.5B
67.2%
146Grok 2 (Dec 12) · proprietary
65.2%
147Mistral Large 3 675B Instruct 2512 · 675B
65.1%
148DeepSeek v3 0324 · 684.5B
64.8%
149GPT 4o (May 13, 2024) · proprietary
64.5%
150Qwen3 14B · 14.8B
64.0%
151Gemini 2.0 Flash 001 · proprietary
63.2%
152Qwen2.5 72B Instruct · 72.7B
62.9%
153Gemini 2.5 Flash Lite Preview Thinking (Jun 17) · proprietary
62.8%
154DeepSeek Chat · proprietary
62.7%
155GPT 4 (Jun 13) · proprietary
62.7%
156GPT 5 Nano (Aug 07, 2025, high) · proprietary
62.7%
157Mistral Medium 2505 · proprietary
62.3%
158Llama 4 Maverick 17B 128E Instruct · 401.6B
61.9%
159Claude 3 Opus (Feb 29, 2024) · proprietary
61.6%
160GPT 4 Turbo · proprietary
61.6%
161Llama 3.1 405B Instruct · 405.9B
61.4%
162C4ai Command A 03 2025 · 111.1B
61.3%
163Magistral Small 2509 · 24.0B
61.3%
164Mistral Large Instruct 2407 · 122.6B
61.2%
165Mistral Large Instruct 2411 · 122.6B
60.8%
166Qwen3 30B A3B · 30.5B
60.3%
167Llama 3.1 70B Instruct · 70.6B
60.0%
168Mistral Small 3.2 24B Instruct 2506 · 24.0B
59.9%
169Qwen3 8B · 8.2B
59.7%
170Llama 3.3 70B Instruct · 70.6B
59.5%
171Gemini 1.5 Pro 002 · proprietary
59.0%
172Mistral Small 3.1 24B Instruct 2503 · 24.0B
58.6%
173Llama 4 Scout 17B 16E Instruct · 108.6B
57.9%
174Claude 3.5 Haiku (Oct 22, 2024) · proprietary
56.7%
175Gemini 1.5 Pro 001 · proprietary
56.5%
176Mistral Large 2402 · proprietary
55.9%
177Mixtral 8x22B Instruct v0.1 · 140.6B
55.1%
178C4ai Command R Plus 08 2024 · 103.8B
54.9%
179GPT 4o Mini (Jul 18, 2024) · proprietary
54.4%
180Meta Llama 3 70B Instruct · 70.6B
54.2%
181Gemini 1.5 Flash 002 · proprietary
53.8%
182Claude 3 Sonnet (Feb 29, 2024) · proprietary
53.6%
183Mistral Medium 2312 · proprietary
53.3%
184Mistral Small 2402 · proprietary
53.0%
185Gemini 2.0 Flash Lite · proprietary
52.5%
186Gemma 3 27B IT · 27.4B
52.5%
187GPT 4.1 Nano (Apr 14, 2025) · proprietary
52.5%
188Claude 2.0 · proprietary
51.8%
189Ministral 3B 2410 · proprietary
51.7%
190Gemini 1.5 Flash 001 · proprietary
51.4%
191Claude 2.1 · proprietary
50.9%
192Gemma 3 4B IT · 4.3B
50.9%
193Llama 3.1 8B Instruct · 8.0B
50.9%
194Claude 3 Haiku (Mar 07, 2024) · proprietary
50.1%
195Gemini 1.5 Flash 8B 001 · proprietary
50.0%
196Mixtral 8x7B Instruct v0.1 · 46.7B
49.6%
197Mistral Small Instruct 2409 · 22.2B
49.5%
198Gemma 3 12B IT · 12.2B
48.8%
199Mistral Nemo Instruct 2407 · 12.2B
48.6%
200GPT 3.5 Turbo (Jan 25) · proprietary
48.5%
201Gemma 2 27B · proprietary
48.0%
202Qwen2.5 7B Instruct · 7.6B
47.7%
203C4ai Command R 08 2024 · 32.3B
46.4%
204Gemini 1.0 Pro 001 · proprietary
45.9%
205Claude Instant 1.2 · proprietary
45.8%
206Ministral 8B 2410 · proprietary
45.7%
207Meta Llama 3 8B Instruct · 8.0B
43.9%
208Open Mistral 7B · proprietary
42.5%
209Llama 2 13B Chat HF · 13.0B
42.2%
210Llama 2 70B Chat HF · 69.0B
41.6%

Score vs model size

Which models give the most quality for their size — the ones worth running locally.

10B100B1Tmodel size (log scale) →93.9%41.6%Kimi K3 · 2.8T · 91.2%Kimi K2.6 · 1T · 90.9%DeepSeek V4 Pro · 1.6T · 90.7%NVIDIA Nemotron 3 Ultra 550B A55B BF16 · 561B · 90.1%DeepSeek V4.1 Flash · 763B · 89.9%DeepSeek V3.2 Exp · 685B · 87.7%GLM 5.3 · 753B · 87.7%Qwen3.5 397B A17B · 403B · 87.5%Inkling · 952B · 87.5%DeepSeek V4 Flash · 291B · 86.4%DeepSeek V3.2 · 685B · 85.6%MiMo V2.5 Pro · 1T · 84.5%Qwen3.5 122B A10B · 125B · 84.3%DeepSeek V3.1 · 685B · 82.7%Gemma 4 31B IT · 31B · 82.7%Qwen3.5 27B · 28B · 82.4%DeepSeek V3.1 Terminus · 685B · 81.3%Qwen3 235B A22B Thinking 2507 · 235B · 80.3%Qwen3.5 35B A3B · 36B · 80.0%MiniMax M3 · 427B · 78.9%Qwen3 235B A22B Instruct 2507 · 235B · 78.4%Qwen3.6 27B · 28B · 78.1%GPT OSS 120B · 117B · 76.3%Qwen3 235B A22B · 235B · 75.7%Qwen3.6 35B A3B · 36B · 73.9%Qwen3 30B A3B Thinking 2507 · 31B · 69.3%GPT OSS 20B · 21B · 68.0%Qwen3 32B · 33B · 67.5%Qwen3 30B A3B Instruct 2507 · 31B · 67.2%Mistral Large 3 675B Instruct 2512 · 675B · 65.1%DeepSeek v3 0324 · 685B · 64.8%Qwen3 14B · 15B · 64.0%Qwen2.5 72B Instruct · 73B · 62.9%Llama 4 Maverick 17B 128E Instruct · 402B · 61.9%Llama 3.1 405B Instruct · 406B · 61.4%C4ai Command A 03 2025 · 111B · 61.3%Magistral Small 2509 · 24B · 61.3%Mistral Large Instruct 2407 · 123B · 61.2%Mistral Large Instruct 2411 · 123B · 60.8%Qwen3 30B A3B · 31B · 60.3%Llama 3.1 70B Instruct · 71B · 60.0%Mistral Small 3.2 24B Instruct 2506 · 24B · 59.9%Llama 3.3 70B Instruct · 71B · 59.5%Mistral Small 3.1 24B Instruct 2503 · 24B · 58.6%Llama 4 Scout 17B 16E Instruct · 109B · 57.9%Mixtral 8x22B Instruct v0.1 · 141B · 55.1%C4ai Command R Plus 08 2024 · 104B · 54.9%Meta Llama 3 70B Instruct · 71B · 54.2%Gemma 3 27B IT · 27B · 52.5%Llama 3.1 8B Instruct · 8B · 50.9%Mixtral 8x7B Instruct v0.1 · 47B · 49.6%Mistral Small Instruct 2409 · 22B · 49.5%Gemma 3 12B IT · 12B · 48.8%Mistral Nemo Instruct 2407 · 12B · 48.6%Qwen2.5 7B Instruct · 8B · 47.7%C4ai Command R 08 2024 · 32B · 46.4%Meta Llama 3 8B Instruct · 8B · 43.9%Llama 2 13B Chat HF · 13B · 42.2%Llama 2 70B Chat HF · 69B · 41.6%Gemma 3 4B IT · 4B · 50.9%Gemma 3 4B ITQwen3 8B · 8B · 59.7%Qwen3 8BQwen3.5 9B · 10B · 71.2%Qwen3.5 9BGemma 4 26B A4B IT · 26B · 74.9%Gemma 4 26B A4B ITQwen3.8 27B · 28B · 88.0%Qwen3.8 27BDeepSeek V4 Flash 0731 · 304B · 90.9%DeepSeek V4 Flash 0731GLM 5.2 · 753B · 93.6%GLM 5.2DeepSeek V4 Pro 0813 · 1.7T · 93.9%DeepSeek V4 Pro 0813
Each dot is a model. Up = higher score, left = smaller (easier to run locally). The dashed line marks the efficiency frontier — the best score you can get at each size or smaller.
  • Gemma 3 4B IT, 4B, score 50.9% — on the efficiency frontier (best score at its size or smaller).
  • Qwen3 8B, 8B, score 59.7% — on the efficiency frontier (best score at its size or smaller).
  • Qwen3.5 9B, 10B, score 71.2% — on the efficiency frontier (best score at its size or smaller).
  • Gemma 4 26B A4B IT, 26B, score 74.9% — on the efficiency frontier (best score at its size or smaller).
  • Qwen3.8 27B, 28B, score 88.0% — on the efficiency frontier (best score at its size or smaller).
  • DeepSeek V4 Flash 0731, 304B, score 90.9% — on the efficiency frontier (best score at its size or smaller).
  • GLM 5.2, 753B, score 93.6% — on the efficiency frontier (best score at its size or smaller).
  • DeepSeek V4 Pro 0813, 1.7T, score 93.9% — on the efficiency frontier (best score at its size or smaller).

DTBench: frequently asked questions

What is the best open LLM on DTBench?
DeepSeek V4 Pro 0813 is the top open model on DTBench, scoring 93.9%. Among all models tested — including proprietary ones — it ranks #44. The top model overall is Claude Fable 5 Max (Anthropic) at 98.4%.
What's the best DTBench model you can run on a 24 GB GPU?
Qwen3.8 27B is the highest-scoring open model that fits in 24 GB at 4-bit quantization (about 15 GB), scoring 88.0% on DTBench.
What's the best DTBench model you can run on a 12 GB GPU?
Qwen3.5 9B is the highest-scoring open model that fits in 12 GB at 4-bit quantization (about 5 GB), scoring 71.2% on DTBench.
Can open models match proprietary models on DTBench?
Not quite on DTBench: the strongest proprietary model (Claude Fable 5 Max) scores 98.4%, ahead of the best open model (DeepSeek V4 Pro 0813) at 93.9% — but you can run the open one yourself.

Scores aggregated from epoch. llmrun does not run this benchmark — see the source for methodology, or the about benchmarks for what it measures.