Reasoning

CritPt Leaderboard

CritPt (Critical Points) tests graduate-level physics problem-solving on research-style questions spanning mechanics, electromagnetism, quantum mechanics, and more. It was created by physicists to probe whether a model can reason through multi-step physics rather than pattern-match textbook answers.

Source: epoch50 open models ranked+117 proprietaryData through Sep 2026

Open models ranked on CritPt

# shows rank among open models / rank overall (including proprietary).

#ModelScore
1 / 27Kimi K3 · 2779.9B
23.4%
2 / 32GLM 5.2 · 753.3B
20.9%
3 / 38GLM 5.3 · 753.3B
19.1%
4 / 41DeepSeek V4 Pro 0813 · 1650.5B
18.0%
5 / 48DeepSeek V4 Flash 0731 · 304.2B
16.6%
6 / 50GLM 5.3 Flash · 321.3B
15.4%
7 / 57DeepSeek V4 Pro · 1598.8B
12.9%
8 / 64Kimi K2.7 Code · 1026.9B
10.0%
9 / 70Inkling Small · 266.0B
8.3%
10 / 73Kimi K2.6 · 1026.9B
8.0%
11 / 74DeepSeek V4 Flash · 290.9B
7.1%
12 / 78Inkling · 952.4B
5.4%
13 / 79Qwen3.8 27B · 27.8B
5.4%
14 / 83GLM 5.1 · 753.9B
4.6%
15 / 85MiMo V2.5 Pro · 1023.2B
4.0%
16 / 86MiMo V2.5 · 310.8B
3.7%
17 / 87MiniMax M3 · 427.0B
3.7%
18 / 88Ring 2.6 1T · 1025.7B
3.7%
19 / 90Kimi K2.5 · 1026.9B
3.1%
20 / 93DeepSeek V3.2 Exp · 685.4B
2.9%
21 / 96Step 3.7 Flash · 201.4B
2.3%
22 / 100GLM 4.7 · 358.3B
1.7%
23 / 101Gemma 4 31B IT · 31.3B
1.4%
24 / 103GPT OSS 20B · 20.9B
1.4%
25 / 108GLM 4.6 · 356.8B
1.1%
26 / 109GPT OSS 120B · 116.8B
1.1%
27 / 110DeepSeek R1 · 684.5B
1.1%
28 / 112Qwen3.5 122B A10B · 125.1B
0.9%
29 / 113Qwen3.6 27B · 27.8B
0.9%
30 / 114Trinity Large Thinking · 398.6B
0.9%
31 / 117MiniMax M2.7 · 228.7B
0.6%
32 / 126Qwen3 30B A3B Thinking 2507 · 30.5B
0.3%
33 / 127Qwen3 32B · 32.8B
0.3%
34 / 128Qwen3.5 9B · 9.7B
0.3%
35 / 129Qwen3.6 35B A3B · 36.0B
0.3%
36 / 133DeepSeek v3 · 684.5B
0.0%
37 / 134DeepSeek v3 0324 · 684.5B
0.0%
38 / 137Gemma 3 12B IT · 12.2B
0.0%
39 / 138Gemma 3 27B IT · 27.4B
0.0%
40 / 146Granite 4.1 30B · 28.9B
0.0%
41 / 149Llama 3.1 8B Instruct · 8.0B
0.0%
42 / 150Llama 3.3 70B Instruct · 70.6B
0.0%
43 / 151Llama 4 Maverick 17B 128E Instruct · 401.6B
0.0%
44 / 152Llama 4 Scout 17B 16E Instruct · 108.6B
0.0%
45 / 153MiMo v2 Flash · 309.8B
0.0%
46 / 162Phi 4 Mini Instruct · 3.8B
0.0%
47 / 163Qwen3 14B · 14.8B
0.0%
48 / 164Qwen3 235B A22B Thinking 2507 · 235.1B
0.0%
49 / 165Qwen3 8B · 8.2B
0.0%
50 / 166Qwen3 Coder Next · 79.7B
0.0%

Score vs model size

Which models give the most quality for their size — the ones worth running locally.

10B100B1Tmodel size (log scale) →23.4%0.0%GLM 5.3 · 753B · 19.1%DeepSeek V4 Pro 0813 · 1.7T · 18.0%GLM 5.3 Flash · 321B · 15.4%DeepSeek V4 Pro · 1.6T · 12.9%Kimi K2.7 Code · 1T · 10.0%Kimi K2.6 · 1T · 8.0%DeepSeek V4 Flash · 291B · 7.1%Inkling · 952B · 5.4%GLM 5.1 · 754B · 4.6%MiMo V2.5 Pro · 1T · 4.0%Ring 2.6 1T · 1T · 3.7%MiniMax M3 · 427B · 3.7%MiMo V2.5 · 311B · 3.7%Kimi K2.5 · 1T · 3.1%DeepSeek V3.2 Exp · 685B · 2.9%Step 3.7 Flash · 201B · 2.3%GLM 4.7 · 358B · 1.7%Gemma 4 31B IT · 31B · 1.4%GPT OSS 120B · 117B · 1.1%GLM 4.6 · 357B · 1.1%DeepSeek R1 · 685B · 1.1%Trinity Large Thinking · 399B · 0.9%Qwen3.5 122B A10B · 125B · 0.9%Qwen3.6 27B · 28B · 0.9%MiniMax M2.7 · 229B · 0.6%Qwen3 30B A3B Thinking 2507 · 31B · 0.3%Qwen3 32B · 33B · 0.3%Qwen3.6 35B A3B · 36B · 0.3%DeepSeek v3 · 685B · 0.0%DeepSeek v3 0324 · 685B · 0.0%Gemma 3 12B IT · 12B · 0.0%Gemma 3 27B IT · 27B · 0.0%Granite 4.1 30B · 29B · 0.0%Llama 3.1 8B Instruct · 8B · 0.0%Llama 3.3 70B Instruct · 71B · 0.0%Llama 4 Maverick 17B 128E Instruct · 402B · 0.0%Llama 4 Scout 17B 16E Instruct · 109B · 0.0%Qwen3 14B · 15B · 0.0%Qwen3 235B A22B Thinking 2507 · 235B · 0.0%Qwen3 8B · 8B · 0.0%Qwen3 Coder Next · 80B · 0.0%MiMo v2 Flash · 310B · 0.0%Phi 4 Mini Instruct · 4B · 0.0%Phi 4 Mini InstructQwen3.5 9B · 10B · 0.3%Qwen3.5 9BGPT OSS 20B · 21B · 1.4%GPT OSS 20BQwen3.8 27B · 28B · 5.4%Qwen3.8 27BInkling Small · 266B · 8.3%Inkling SmallDeepSeek V4 Flash 0731 · 304B · 16.6%DeepSeek V4 Flash 0731GLM 5.2 · 753B · 20.9%GLM 5.2Kimi K3 · 2.8T · 23.4%Kimi K3
Each dot is a model. Up = higher score, left = smaller (easier to run locally). The dashed line marks the efficiency frontier — the best score you can get at each size or smaller.
  • Phi 4 Mini Instruct, 4B, score 0.0% — on the efficiency frontier (best score at its size or smaller).
  • Qwen3.5 9B, 10B, score 0.3% — on the efficiency frontier (best score at its size or smaller).
  • GPT OSS 20B, 21B, score 1.4% — on the efficiency frontier (best score at its size or smaller).
  • Qwen3.8 27B, 28B, score 5.4% — on the efficiency frontier (best score at its size or smaller).
  • Inkling Small, 266B, score 8.3% — on the efficiency frontier (best score at its size or smaller).
  • DeepSeek V4 Flash 0731, 304B, score 16.6% — on the efficiency frontier (best score at its size or smaller).
  • GLM 5.2, 753B, score 20.9% — on the efficiency frontier (best score at its size or smaller).
  • Kimi K3, 2.8T, score 23.4% — on the efficiency frontier (best score at its size or smaller).

CritPt: frequently asked questions

What is the best open LLM on CritPt?
Kimi K3 is the top open model on CritPt, scoring 23.4%. Among all models tested — including proprietary ones — it ranks #27. The top model overall is GPT 5.6 Sol Max (OpenAI) at 32.3%.
What's the best CritPt model you can run on a 24 GB GPU?
Qwen3.8 27B is the highest-scoring open model that fits in 24 GB at 4-bit quantization (about 15 GB), scoring 5.4% on CritPt.
What's the best CritPt model you can run on a 12 GB GPU?
GPT OSS 20B is the highest-scoring open model that fits in 12 GB at 4-bit quantization (about 12 GB), scoring 1.4% on CritPt.
Can open models match proprietary models on CritPt?
Not quite on CritPt: the strongest proprietary model (GPT 5.6 Sol Max) scores 32.3%, ahead of the best open model (Kimi K3) at 23.4% — but you can run the open one yourself.

Scores aggregated from epoch. llmrun does not run this benchmark — see the source for methodology, or the about benchmarks for what it measures.