Reasoning

CritPt Leaderboard

CritPt (Critical Points) tests graduate-level physics problem-solving on research-style questions spanning mechanics, electromagnetism, quantum mechanics, and more. It was created by physicists to probe whether a model can reason through multi-step physics rather than pattern-match textbook answers.

Source: epoch50 open models ranked+117 proprietaryData through Sep 2026

All models ranked on CritPt

Proprietary / closed models are shown dimmed — you can't run them locally, but they show where the open field stands.

#ModelScore
1GPT 5.6 Sol Max · proprietary
32.3%
2GPT 6 Astra Max · proprietary
31.7%
3GPT 6 Astra (xhigh) · proprietary
31.4%
4Claude Fable 5.1 (xhigh) · proprietary
31.1%
5GPT 5.5 Pro (xhigh) · proprietary
30.6%
6Claude Fable 5.1 (high) · proprietary
30.3%
7GPT 5.4 Pro (Mar 05, 2026, xhigh) · proprietary
30.0%
8GPT 5.6 Terra Max · proprietary
30.0%
9Claude Fable 5.1 Max · proprietary
29.7%
10Claude Fable 5.1 (medium) · proprietary
29.1%
11Claude Opus 5 Max · proprietary
29.1%
12GPT 6 Astra (medium) · proprietary
29.1%
13GPT 6 Astra (high) · proprietary
28.9%
14GPT 5.6 Sol (xhigh) · proprietary
28.6%
15Claude Fable 5 Max · proprietary
28.6%
16Claude Opus 5 (high) · proprietary
28.3%
17Claude Fable 5.1 (low) · proprietary
27.7%
18Claude Opus 5 (xhigh) · proprietary
27.7%
19GPT 5.5 (xhigh) · proprietary
27.1%
20GPT 5.6 Terra (xhigh) · proprietary
27.1%
21Claude Opus 5 (medium) · proprietary
26.9%
22GPT 6 Astra (low) · proprietary
26.3%
23Gemini 3 Deep Think Preview · proprietary
25.7%
24GPT 5.6 Sol (high) · proprietary
25.7%
25GPT 5.5 (high) · proprietary
25.4%
26GPT 5.4 (Mar 05, 2026, xhigh) · proprietary
23.4%
27Kimi K3 · 2779.9B
23.4%
28Claude Opus 5 (low) · proprietary
23.1%
29GPT 5.6 Sol (medium) · proprietary
22.9%
30GPT 5.6 Terra (high) · proprietary
22.9%
31Claude Opus 4.8 Max · proprietary
20.9%
32GLM 5.2 · 753.3B
20.9%
33GPT 5.6 Luna Max · proprietary
20.6%
34GPT 5.6 Luna (xhigh) · proprietary
20.6%
35Qwen3.8 Max (unspecified) · proprietary
20.0%
36GPT 6 Astra None · proprietary
19.7%
37Grok 4.6 (xhigh) · proprietary
19.7%
38GLM 5.3 · 753.3B
19.1%
39GPT 5.5 (medium) · proprietary
18.6%
40Gemini 3.8 Flash (high) · proprietary
18.3%
41DeepSeek V4 Pro 0813 · 1650.5B
18.0%
42Gemini 3.1 Pro Preview · proprietary
17.7%
43Grok 4.6 (medium) · proprietary
17.7%
44Muse Spark 1.2 (xhigh) · proprietary
17.7%
45GPT 5.6 Terra (medium) · proprietary
17.4%
46Grok 4.6 (high) · proprietary
17.1%
47Claude Sonnet 5 Max · proprietary
16.9%
48DeepSeek V4 Flash 0731 · 304.2B
16.6%
49GPT 5.6 Luna (high) · proprietary
16.6%
50GLM 5.3 Flash · 321.3B
15.4%
51Grok 4.5 (high) · proprietary
15.4%
52Muse Spark 1.1 · proprietary
15.1%
53GPT 5.6 Sol (low) · proprietary
14.9%
54Gemini 3.7 Flash (high) · proprietary
14.3%
55Qwen3.7 Max · proprietary
13.4%
56Gemini 3.5 Flash (high) · proprietary
13.1%
57DeepSeek V4 Pro · 1598.8B
12.9%
58GPT 5 (Aug 07, 2025, high) · proprietary
12.6%
59Gemini 3.8 Flash (medium) · proprietary
12.3%
60Claude Opus 4.7 Max · proprietary
12.0%
61Muse Spark · proprietary
11.3%
62Gemini 3.6 Flash (high) · proprietary
10.6%
63GPT 5.4 Mini (Mar 17, 2026, xhigh) · proprietary
10.0%
64Kimi K2.7 Code · 1026.9B
10.0%
65Gemini 3.7 Flash (medium) · proprietary
9.4%
66GPT 5.6 Terra (low) · proprietary
9.4%
67GPT 5.4 Nano (Mar 17, 2026, xhigh) · proprietary
9.3%
68Grok Build 0.1 · proprietary
9.1%
69Qwen3.7 Plus · proprietary
9.1%
70Inkling Small · 266.0B
8.3%
71GPT 5.5 (low) · proprietary
8.0%
72Grok 4.3 (high) · proprietary
8.0%
73Kimi K2.6 · 1026.9B
8.0%
74DeepSeek V4 Flash · 290.9B
7.1%
75Gemini 3 Pro Preview · proprietary
6.9%
76Gemini 3.7 Flash (low) · proprietary
5.7%
77Grok 4.6 (low) · proprietary
5.7%
78Inkling · 952.4B
5.4%
79Qwen3.8 27B · 27.8B
5.4%
80GPT 5.6 Sol None · proprietary
5.1%
81GPT 5.1 (Nov 13, 2025, unspecified) · proprietary
4.9%
82GPT 5.6 Luna (medium) · proprietary
4.9%
83GLM 5.1 · 753.9B
4.6%
84Gemini 3.8 Flash (low) · proprietary
4.0%
85MiMo V2.5 Pro · 1023.2B
4.0%
86MiMo V2.5 · 310.8B
3.7%
87MiniMax M3 · 427.0B
3.7%
88Ring 2.6 1T · 1025.7B
3.7%
89Claude Sonnet 4.6 Max · proprietary
3.1%
90Kimi K2.5 · 1026.9B
3.1%
91Nemotron 3 Super · proprietary
3.1%
92Nemotron 3 Ultra · proprietary
3.1%
93DeepSeek V3.2 Exp · 685.4B
2.9%
94Qwen3.6 Plus · proprietary
2.9%
95GPT 5.6 Luna (low) · proprietary
2.6%
96Step 3.7 Flash · 201.4B
2.3%
97Gemini 2.5 Pro · proprietary
2.0%
98GPT 5.6 Terra None · proprietary
2.0%
99DeepSeek V3.1 Terminus · proprietary
1.7%
100GLM 4.7 · 358.3B
1.7%
101Gemma 4 31B IT · 31.3B
1.4%
102GPT 5.5 None · proprietary
1.4%
103GPT OSS 20B · 20.9B
1.4%
104O3 (Apr 16, 2025, high) · proprietary
1.4%
105Claude Sonnet 4.5 (Sep 29, 2025, unspecified) · proprietary
1.1%
106Claude Sonnet 5 (high) · proprietary
1.1%
107Gemini 3.1 Flash Lite · proprietary
1.1%
108GLM 4.6 · 356.8B
1.1%
109GPT OSS 120B · 116.8B
1.1%
110DeepSeek R1 · 684.5B
1.1%
111Gemini 2.5 Flash · proprietary
1.1%
112Qwen3.5 122B A10B · 125.1B
0.9%
113Qwen3.6 27B · 27.8B
0.9%
114Trinity Large Thinking · 398.6B
0.9%
115Mercury 2 · proprietary
0.9%
116O4 Mini (Apr 16, 2025, high) · proprietary
0.6%
117MiniMax M2.7 · 228.7B
0.6%
118Qwen3.5 35B A3B None · proprietary
0.6%
119Claude Opus 4 (May 14, 2025, unspecified) · proprietary
0.3%
120Claude Sonnet 4 (May 14, 2025, unspecified) · proprietary
0.3%
121Command A Plus (May 2026) · proprietary
0.3%
122GPT 5.6 Luna None · proprietary
0.3%
123Magistral Medium 2509 · proprietary
0.3%
124Magistral Small 2509 · proprietary
0.3%
125O3 Mini (Jan 31, 2025, high) · proprietary
0.3%
126Qwen3 30B A3B Thinking 2507 · 30.5B
0.3%
127Qwen3 32B · 32.8B
0.3%
128Qwen3.5 9B · 9.7B
0.3%
129Qwen3.6 35B A3B · 36.0B
0.3%
130Claude 3.5 Haiku (Oct 22, 2024) · proprietary
0.0%
131Claude Haiku 4.5 (Oct 01, 2025, unspecified) · proprietary
0.0%
132Cogito 671B v2.1 · proprietary
0.0%
133DeepSeek v3 · 684.5B
0.0%
134DeepSeek v3 0324 · 684.5B
0.0%
135Devstral Small 2512 · proprietary
0.0%
136Gemini 3.5 Flash Lite · proprietary
0.0%
137Gemma 3 12B IT · 12.2B
0.0%
138Gemma 3 27B IT · 27.4B
0.0%
139Gemma 4 26B A4B · proprietary
0.0%
140GPT 4.1 Mini (Apr 14, 2025) · proprietary
0.0%
141GPT 4.1 Nano (Apr 14, 2025) · proprietary
0.0%
142GPT 4o (Nov 20, 2024) · proprietary
0.0%
143GPT 5 (Aug 07, 2025, minimal) · proprietary
0.0%
144GPT 5 Mini (Aug 07, 2025, unspecified) · proprietary
0.0%
145GPT 5.5 Instant · proprietary
0.0%
146Granite 4.1 30B · 28.9B
0.0%
147Grok 4.3 (unspecified) · proprietary
0.0%
148Grok 4.3 None · proprietary
0.0%
149Llama 3.1 8B Instruct · 8.0B
0.0%
150Llama 3.3 70B Instruct · 70.6B
0.0%
151Llama 4 Maverick 17B 128E Instruct · 401.6B
0.0%
152Llama 4 Scout 17B 16E Instruct · 108.6B
0.0%
153MiMo v2 Flash · 309.8B
0.0%
154Mistral Large 2512 · proprietary
0.0%
155Mistral Medium 2508 · proprietary
0.0%
156Mistral Medium 2604 · proprietary
0.0%
157Mistral Small 2503 · proprietary
0.0%
158Mistral Small 2506 · proprietary
0.0%
159Nova 2.0 Pro Preview (low) · proprietary
0.0%
160Nova 2.0 Pro Preview (medium) · proprietary
0.0%
161Nova 2.0 Pro Preview None · proprietary
0.0%
162Phi 4 Mini Instruct · 3.8B
0.0%
163Qwen3 14B · 14.8B
0.0%
164Qwen3 235B A22B Thinking 2507 · 235.1B
0.0%
165Qwen3 8B · 8.2B
0.0%
166Qwen3 Coder Next · 79.7B
0.0%
167Solar Pro 3 · proprietary
0.0%

Score vs model size

Which models give the most quality for their size — the ones worth running locally.

10B100B1Tmodel size (log scale) →23.4%0.0%GLM 5.3 · 753B · 19.1%DeepSeek V4 Pro 0813 · 1.7T · 18.0%GLM 5.3 Flash · 321B · 15.4%DeepSeek V4 Pro · 1.6T · 12.9%Kimi K2.7 Code · 1T · 10.0%Kimi K2.6 · 1T · 8.0%DeepSeek V4 Flash · 291B · 7.1%Inkling · 952B · 5.4%GLM 5.1 · 754B · 4.6%MiMo V2.5 Pro · 1T · 4.0%Ring 2.6 1T · 1T · 3.7%MiniMax M3 · 427B · 3.7%MiMo V2.5 · 311B · 3.7%Kimi K2.5 · 1T · 3.1%DeepSeek V3.2 Exp · 685B · 2.9%Step 3.7 Flash · 201B · 2.3%GLM 4.7 · 358B · 1.7%Gemma 4 31B IT · 31B · 1.4%GPT OSS 120B · 117B · 1.1%GLM 4.6 · 357B · 1.1%DeepSeek R1 · 685B · 1.1%Trinity Large Thinking · 399B · 0.9%Qwen3.5 122B A10B · 125B · 0.9%Qwen3.6 27B · 28B · 0.9%MiniMax M2.7 · 229B · 0.6%Qwen3 30B A3B Thinking 2507 · 31B · 0.3%Qwen3 32B · 33B · 0.3%Qwen3.6 35B A3B · 36B · 0.3%DeepSeek v3 · 685B · 0.0%DeepSeek v3 0324 · 685B · 0.0%Gemma 3 12B IT · 12B · 0.0%Gemma 3 27B IT · 27B · 0.0%Granite 4.1 30B · 29B · 0.0%Llama 3.1 8B Instruct · 8B · 0.0%Llama 3.3 70B Instruct · 71B · 0.0%Llama 4 Maverick 17B 128E Instruct · 402B · 0.0%Llama 4 Scout 17B 16E Instruct · 109B · 0.0%Qwen3 14B · 15B · 0.0%Qwen3 235B A22B Thinking 2507 · 235B · 0.0%Qwen3 8B · 8B · 0.0%Qwen3 Coder Next · 80B · 0.0%MiMo v2 Flash · 310B · 0.0%Phi 4 Mini Instruct · 4B · 0.0%Phi 4 Mini InstructQwen3.5 9B · 10B · 0.3%Qwen3.5 9BGPT OSS 20B · 21B · 1.4%GPT OSS 20BQwen3.8 27B · 28B · 5.4%Qwen3.8 27BInkling Small · 266B · 8.3%Inkling SmallDeepSeek V4 Flash 0731 · 304B · 16.6%DeepSeek V4 Flash 0731GLM 5.2 · 753B · 20.9%GLM 5.2Kimi K3 · 2.8T · 23.4%Kimi K3
Each dot is a model. Up = higher score, left = smaller (easier to run locally). The dashed line marks the efficiency frontier — the best score you can get at each size or smaller.
  • Phi 4 Mini Instruct, 4B, score 0.0% — on the efficiency frontier (best score at its size or smaller).
  • Qwen3.5 9B, 10B, score 0.3% — on the efficiency frontier (best score at its size or smaller).
  • GPT OSS 20B, 21B, score 1.4% — on the efficiency frontier (best score at its size or smaller).
  • Qwen3.8 27B, 28B, score 5.4% — on the efficiency frontier (best score at its size or smaller).
  • Inkling Small, 266B, score 8.3% — on the efficiency frontier (best score at its size or smaller).
  • DeepSeek V4 Flash 0731, 304B, score 16.6% — on the efficiency frontier (best score at its size or smaller).
  • GLM 5.2, 753B, score 20.9% — on the efficiency frontier (best score at its size or smaller).
  • Kimi K3, 2.8T, score 23.4% — on the efficiency frontier (best score at its size or smaller).

CritPt: frequently asked questions

What is the best open LLM on CritPt?
Kimi K3 is the top open model on CritPt, scoring 23.4%. Among all models tested — including proprietary ones — it ranks #27. The top model overall is GPT 5.6 Sol Max (OpenAI) at 32.3%.
What's the best CritPt model you can run on a 24 GB GPU?
Qwen3.8 27B is the highest-scoring open model that fits in 24 GB at 4-bit quantization (about 15 GB), scoring 5.4% on CritPt.
What's the best CritPt model you can run on a 12 GB GPU?
GPT OSS 20B is the highest-scoring open model that fits in 12 GB at 4-bit quantization (about 12 GB), scoring 1.4% on CritPt.
Can open models match proprietary models on CritPt?
Not quite on CritPt: the strongest proprietary model (GPT 5.6 Sol Max) scores 32.3%, ahead of the best open model (Kimi K3) at 23.4% — but you can run the open one yourself.

Scores aggregated from epoch. llmrun does not run this benchmark — see the source for methodology, or the about benchmarks for what it measures.