Coding

SciCode Leaderboard

SciCode asks a model to write real scientific-computing code — implementing multi-step numerical methods drawn from physics, chemistry, biology, and math research — and checks the output against reference solutions. It was built by a team of scientists and ML researchers to test coding ability on genuine research problems, not typical software tasks.

Source: epoch50 open models ranked+105 proprietaryData through Sep 2026

All models ranked on SciCode

Proprietary / closed models are shown dimmed — you can't run them locally, but they show where the open field stands.

#ModelScore
1Claude Fable 5.1 Max · proprietary
62.0%
2Claude Fable 5 Max · proprietary
60.2%
3Claude Fable 5.1 (xhigh) · proprietary
60.1%
4Gemini 3.1 Pro Preview · proprietary
58.9%
5Kimi K3 · 2779.9B
58.7%
6Muse Spark 1.1 · proprietary
58.2%
7Gemini 3.7 Flash (medium) · proprietary
57.9%
8Claude Fable 5.1 (high) · proprietary
57.6%
9GPT 5.6 Sol (high) · proprietary
56.9%
10Gemini 3.7 Flash (high) · proprietary
56.8%
11GPT 5.4 (Mar 05, 2026, xhigh) · proprietary
56.6%
12GLM 5.3 · 753.3B
56.5%
13GPT 5.6 Sol (medium) · proprietary
56.5%
14GPT 6 Astra Max · proprietary
56.5%
15Muse Spark 1.2 (xhigh) · proprietary
56.4%
16GPT 5.5 (xhigh) · proprietary
56.1%
17GPT 5.6 Sol Max · proprietary
56.1%
18GPT 5.6 Sol (xhigh) · proprietary
56.0%
19GPT 5.5 (high) · proprietary
55.9%
20Claude Fable 5.1 (low) · proprietary
55.7%
21Claude Opus 5 Max · proprietary
55.7%
22GPT 6 Astra (xhigh) · proprietary
55.7%
23GPT 5.6 Sol (low) · proprietary
55.4%
24GPT 6 Astra (high) · proprietary
55.4%
25Claude Fable 5.1 (medium) · proprietary
55.3%
26Claude Opus 5 (xhigh) · proprietary
55.0%
27Grok 4.6 (medium) · proprietary
54.6%
28Claude Opus 4.7 Max · proprietary
54.5%
29Gemini 3.8 Flash (medium) · proprietary
54.4%
30Claude Opus 5 (high) · proprietary
54.3%
31Gemini 3.8 Flash (low) · proprietary
54.3%
32GPT 6 Astra (medium) · proprietary
54.2%
33GPT 6 Astra (low) · proprietary
54.0%
34Grok 4.5 (high) · proprietary
54.0%
35GPT 5.6 Terra Max · proprietary
53.9%
36Claude Sonnet 5 Max · proprietary
53.6%
37Gemini 3.7 Flash (low) · proprietary
53.6%
38Gemini 3.8 Flash (high) · proprietary
53.6%
39Grok 4.6 (high) · proprietary
53.6%
40Claude Opus 4.8 Max · proprietary
53.5%
41GPT 5.5 (medium) · proprietary
53.5%
42GPT 6 Astra None · proprietary
53.5%
43Kimi K2.6 · 1026.9B
53.5%
44Gemini 3.5 Flash (high) · proprietary
53.1%
45Qwen3.8 Max (unspecified) · proprietary
52.9%
46Gemini 3.6 Flash (high) · proprietary
52.7%
47GPT 5.6 Luna Max · proprietary
52.5%
48GPT 5.5 (low) · proprietary
51.6%
49GPT 5.6 Terra (xhigh) · proprietary
51.6%
50Grok 4.6 (xhigh) · proprietary
51.6%
51Muse Spark · proprietary
51.5%
52Claude Opus 5 (medium) · proprietary
50.7%
53GPT 5.6 Luna (high) · proprietary
50.7%
54GLM 5.2 · 753.3B
50.5%
55Grok Build 0.1 · proprietary
50.2%
56MiMo V2.5 Pro · 1023.2B
50.2%
57GPT 5.6 Terra (high) · proprietary
50.1%
58DeepSeek V4 Pro · 1598.8B
50.0%
59GPT 5.6 Luna (xhigh) · proprietary
50.0%
60DeepSeek V4 Flash 0731 · 304.2B
49.9%
61GPT 5.4 Mini (Mar 17, 2026, xhigh) · proprietary
49.9%
62GPT 5.6 Terra (medium) · proprietary
49.6%
63DeepSeek V4 Pro 0813 · 1650.5B
49.2%
64GPT 5.6 Terra (low) · proprietary
49.2%
65Kimi K2.5 · 1026.9B
49.0%
66Qwen3.7 Max · proprietary
48.8%
67Inkling Small · 266.0B
48.7%
68Claude Sonnet 5 (high) · proprietary
48.6%
69GPT 5.5 Instant · proprietary
48.6%
70Grok 4.6 (low) · proprietary
48.4%
71Claude Opus 5 (low) · proprietary
48.0%
72Kimi K2.7 Code · 1026.9B
47.4%
73GPT 5.5 None · proprietary
47.3%
74Grok 4.3 (high) · proprietary
47.3%
75GPT 5.6 Sol None · proprietary
47.1%
76MiniMax M2.7 · 228.7B
47.0%
77GPT 5.4 Nano (Mar 17, 2026, xhigh) · proprietary
46.9%
78Claude Sonnet 4.6 Max · proprietary
46.8%
79GLM 5.3 Flash · 321.3B
46.1%
80Inkling · 952.4B
46.1%
81GPT 5.6 Luna (medium) · proprietary
45.8%
82GPT 5.6 Luna (low) · proprietary
45.6%
83Qwen3.7 Plus · proprietary
45.5%
84MiniMax M3 · 427.0B
45.4%
85GLM 4.7 · 358.3B
45.1%
86DeepSeek V4 Flash · 290.9B
44.9%
87Claude Sonnet 4.5 (Sep 29, 2025, unspecified) · proprietary
44.7%
88Qwen3.8 27B · 27.8B
44.7%
89GPT 5.6 Terra None · proprietary
44.6%
90GLM 5.1 · 753.9B
43.8%
91Gemma 4 31B IT · 31.3B
43.4%
92Claude Haiku 4.5 (Oct 01, 2025, unspecified) · proprietary
43.3%
93GPT 5.1 (Nov 13, 2025, unspecified) · proprietary
43.3%
94MiMo V2.5 · 310.8B
43.1%
95GPT 5 (Aug 07, 2025, unspecified) · proprietary
42.9%
96Gemini 2.5 Pro · proprietary
42.8%
97Nova 2.0 Pro Preview (medium) · proprietary
42.7%
98Qwen3 235B A22B Thinking 2507 · 235.1B
42.4%
99Ring 2.6 1T · 1025.7B
42.4%
100Gemini 3.1 Flash Lite · proprietary
41.9%
101Cogito 671B v2.1 · proprietary
41.0%
102Gemini 3.5 Flash Lite · proprietary
40.9%
103Qwen3.6 Plus · proprietary
40.7%
104DeepSeek V3.1 Terminus · proprietary
40.6%
105GPT 4.1 Mini (Apr 14, 2025) · proprietary
40.4%
106Claude Sonnet 4 (May 14, 2025, unspecified) · proprietary
40.1%
107Gemma 4 26B A4B · proprietary
40.1%
108Step 3.7 Flash · 201.4B
40.1%
109GPT 5.6 Luna None · proprietary
39.9%
110Nemotron 3 Ultra · proprietary
39.9%
111O3 Mini (Jan 31, 2025, high) · proprietary
39.8%
112Mistral Medium 2604 · proprietary
39.6%
113GPT 5 Mini (Aug 07, 2025, unspecified) · proprietary
39.2%
114Magistral Medium 2509 · proprietary
39.2%
115DeepSeek V3.2 Exp · 685.4B
38.9%
116GPT OSS 120B · 116.8B
38.9%
117Mercury 2 · proprietary
38.7%
118Nova 2.0 Pro Preview (low) · proprietary
38.7%
119GLM 4.6 · 356.8B
38.4%
120Command A Plus (May 2026) · proprietary
37.9%
121Qwen3.6 27B · 27.8B
37.3%
122Mistral Large 2512 · proprietary
36.2%
123Trinity Large Thinking · 398.6B
36.1%
124Nemotron 3 Super · proprietary
36.0%
125DeepSeek v3 0324 · 684.5B
35.8%
126Qwen3.6 35B A3B · 36.0B
35.8%
127DeepSeek R1 · 684.5B
35.7%
128Qwen3.5 122B A10B · 125.1B
35.6%
129DeepSeek v3 · 684.5B
35.4%
130Qwen3 32B · 32.8B
35.4%
131Magistral Small 2509 · proprietary
35.2%
132GPT OSS 20B · 20.9B
34.4%
133Mistral Medium 2508 · proprietary
33.8%
134Qwen3 30B A3B Thinking 2507 · 30.5B
33.3%
135Llama 4 Maverick 17B 128E Instruct · 401.6B
33.1%
136Qwen3 Coder Next · 79.7B
32.3%
137Qwen3 14B · 14.8B
31.6%
138Qwen3.5 35B A3B None · proprietary
29.3%
139Devstral Small 2512 · proprietary
28.8%
140Nova 2.0 Pro Preview None · proprietary
28.1%
141Qwen3.5 9B · 9.7B
27.6%
142Claude 3.5 Haiku (Oct 22, 2024) · proprietary
27.4%
143Mistral Small 2503 · proprietary
26.5%
144Mistral Small 2506 · proprietary
26.4%
145Llama 3.3 70B Instruct · 70.6B
26.0%
146GPT 4.1 Nano (Apr 14, 2025) · proprietary
25.9%
147MiMo v2 Flash · 309.8B
25.9%
148Granite 4.1 30B · 28.9B
25.8%
149Solar Pro 3 · proprietary
24.6%
150Qwen3 8B · 8.2B
22.6%
151Gemma 3 27B IT · 27.4B
21.2%
152Gemma 3 12B IT · 12.2B
17.4%
153Llama 4 Scout 17B 16E Instruct · 108.6B
17.0%
154Llama 3.1 8B Instruct · 8.0B
13.2%
155Phi 4 Mini Instruct · 3.8B
10.8%

Score vs model size

Which models give the most quality for their size — the ones worth running locally.

10B100B1Tmodel size (log scale) →58.7%10.8%Kimi K2.6 · 1T · 53.5%GLM 5.2 · 753B · 50.5%MiMo V2.5 Pro · 1T · 50.2%DeepSeek V4 Pro · 1.6T · 50.0%DeepSeek V4 Pro 0813 · 1.7T · 49.2%Kimi K2.5 · 1T · 49.0%Kimi K2.7 Code · 1T · 47.4%Inkling · 952B · 46.1%GLM 5.3 Flash · 321B · 46.1%MiniMax M3 · 427B · 45.4%GLM 4.7 · 358B · 45.1%DeepSeek V4 Flash · 291B · 44.9%GLM 5.1 · 754B · 43.8%Gemma 4 31B IT · 31B · 43.4%MiMo V2.5 · 311B · 43.1%Ring 2.6 1T · 1T · 42.4%Qwen3 235B A22B Thinking 2507 · 235B · 42.4%Step 3.7 Flash · 201B · 40.1%DeepSeek V3.2 Exp · 685B · 38.9%GPT OSS 120B · 117B · 38.9%GLM 4.6 · 357B · 38.4%Qwen3.6 27B · 28B · 37.3%Trinity Large Thinking · 399B · 36.1%DeepSeek v3 0324 · 685B · 35.8%Qwen3.6 35B A3B · 36B · 35.8%DeepSeek R1 · 685B · 35.7%Qwen3.5 122B A10B · 125B · 35.6%DeepSeek v3 · 685B · 35.4%Qwen3 32B · 33B · 35.4%Qwen3 30B A3B Thinking 2507 · 31B · 33.3%Llama 4 Maverick 17B 128E Instruct · 402B · 33.1%Qwen3 Coder Next · 80B · 32.3%Llama 3.3 70B Instruct · 71B · 26.0%MiMo v2 Flash · 310B · 25.9%Granite 4.1 30B · 29B · 25.8%Gemma 3 27B IT · 27B · 21.2%Gemma 3 12B IT · 12B · 17.4%Llama 4 Scout 17B 16E Instruct · 109B · 17.0%Phi 4 Mini Instruct · 4B · 10.8%Phi 4 Mini InstructLlama 3.1 8B Instruct · 8B · 13.2%Llama 3.1 8B InstructQwen3 8B · 8B · 22.6%Qwen3 8BQwen3.5 9B · 10B · 27.6%Qwen3.5 9BQwen3 14B · 15B · 31.6%Qwen3 14BGPT OSS 20B · 21B · 34.4%GPT OSS 20BQwen3.8 27B · 28B · 44.7%Qwen3.8 27BMiniMax M2.7 · 229B · 47.0%Inkling Small · 266B · 48.7%Inkling SmallDeepSeek V4 Flash 0731 · 304B · 49.9%DeepSeek V4 Flash 0731GLM 5.3 · 753B · 56.5%GLM 5.3Kimi K3 · 2.8T · 58.7%Kimi K3
Each dot is a model. Up = higher score, left = smaller (easier to run locally). The dashed line marks the efficiency frontier — the best score you can get at each size or smaller.
  • Phi 4 Mini Instruct, 4B, score 10.8% — on the efficiency frontier (best score at its size or smaller).
  • Llama 3.1 8B Instruct, 8B, score 13.2% — on the efficiency frontier (best score at its size or smaller).
  • Qwen3 8B, 8B, score 22.6% — on the efficiency frontier (best score at its size or smaller).
  • Qwen3.5 9B, 10B, score 27.6% — on the efficiency frontier (best score at its size or smaller).
  • Qwen3 14B, 15B, score 31.6% — on the efficiency frontier (best score at its size or smaller).
  • GPT OSS 20B, 21B, score 34.4% — on the efficiency frontier (best score at its size or smaller).
  • Qwen3.8 27B, 28B, score 44.7% — on the efficiency frontier (best score at its size or smaller).
  • MiniMax M2.7, 229B, score 47.0% — on the efficiency frontier (best score at its size or smaller).
  • Inkling Small, 266B, score 48.7% — on the efficiency frontier (best score at its size or smaller).
  • DeepSeek V4 Flash 0731, 304B, score 49.9% — on the efficiency frontier (best score at its size or smaller).
  • GLM 5.3, 753B, score 56.5% — on the efficiency frontier (best score at its size or smaller).
  • Kimi K3, 2.8T, score 58.7% — on the efficiency frontier (best score at its size or smaller).

SciCode: frequently asked questions

What is the best open LLM on SciCode?
Kimi K3 is the top open model on SciCode, scoring 58.7%. Among all models tested — including proprietary ones — it ranks #5. The top model overall is Claude Fable 5.1 Max (Anthropic) at 62.0%.
What's the best SciCode model you can run on a 24 GB GPU?
Qwen3.8 27B is the highest-scoring open model that fits in 24 GB at 4-bit quantization (about 15 GB), scoring 44.7% on SciCode.
What's the best SciCode model you can run on a 12 GB GPU?
GPT OSS 20B is the highest-scoring open model that fits in 12 GB at 4-bit quantization (about 12 GB), scoring 34.4% on SciCode.
Can open models match proprietary models on SciCode?
Not quite on SciCode: the strongest proprietary model (Claude Fable 5.1 Max) scores 62.0%, ahead of the best open model (Kimi K3) at 58.7% — but you can run the open one yourself.

Scores aggregated from epoch. llmrun does not run this benchmark — see the source for methodology, or the about benchmarks for what it measures.