Coding

SciCode Leaderboard

SciCode asks a model to write real scientific-computing code — implementing multi-step numerical methods drawn from physics, chemistry, biology, and math research — and checks the output against reference solutions. It was built by a team of scientists and ML researchers to test coding ability on genuine research problems, not typical software tasks.

Source: epoch50 open models ranked+105 proprietaryData through Sep 2026

Open models ranked on SciCode

# shows rank among open models / rank overall (including proprietary).

#ModelScore
1 / 5Kimi K3 · 2779.9B
58.7%
2 / 12GLM 5.3 · 753.3B
56.5%
3 / 43Kimi K2.6 · 1026.9B
53.5%
4 / 54GLM 5.2 · 753.3B
50.5%
5 / 56MiMo V2.5 Pro · 1023.2B
50.2%
6 / 58DeepSeek V4 Pro · 1598.8B
50.0%
7 / 60DeepSeek V4 Flash 0731 · 304.2B
49.9%
8 / 63DeepSeek V4 Pro 0813 · 1650.5B
49.2%
9 / 65Kimi K2.5 · 1026.9B
49.0%
10 / 67Inkling Small · 266.0B
48.7%
11 / 72Kimi K2.7 Code · 1026.9B
47.4%
12 / 76MiniMax M2.7 · 228.7B
47.0%
13 / 79GLM 5.3 Flash · 321.3B
46.1%
14 / 80Inkling · 952.4B
46.1%
15 / 84MiniMax M3 · 427.0B
45.4%
16 / 85GLM 4.7 · 358.3B
45.1%
17 / 86DeepSeek V4 Flash · 290.9B
44.9%
18 / 88Qwen3.8 27B · 27.8B
44.7%
19 / 90GLM 5.1 · 753.9B
43.8%
20 / 91Gemma 4 31B IT · 31.3B
43.4%
21 / 94MiMo V2.5 · 310.8B
43.1%
22 / 98Qwen3 235B A22B Thinking 2507 · 235.1B
42.4%
23 / 99Ring 2.6 1T · 1025.7B
42.4%
24 / 108Step 3.7 Flash · 201.4B
40.1%
25 / 115DeepSeek V3.2 Exp · 685.4B
38.9%
26 / 116GPT OSS 120B · 116.8B
38.9%
27 / 119GLM 4.6 · 356.8B
38.4%
28 / 121Qwen3.6 27B · 27.8B
37.3%
29 / 123Trinity Large Thinking · 398.6B
36.1%
30 / 125DeepSeek v3 0324 · 684.5B
35.8%
31 / 126Qwen3.6 35B A3B · 36.0B
35.8%
32 / 127DeepSeek R1 · 684.5B
35.7%
33 / 128Qwen3.5 122B A10B · 125.1B
35.6%
34 / 129DeepSeek v3 · 684.5B
35.4%
35 / 130Qwen3 32B · 32.8B
35.4%
36 / 132GPT OSS 20B · 20.9B
34.4%
37 / 134Qwen3 30B A3B Thinking 2507 · 30.5B
33.3%
38 / 135Llama 4 Maverick 17B 128E Instruct · 401.6B
33.1%
39 / 136Qwen3 Coder Next · 79.7B
32.3%
40 / 137Qwen3 14B · 14.8B
31.6%
41 / 141Qwen3.5 9B · 9.7B
27.6%
42 / 145Llama 3.3 70B Instruct · 70.6B
26.0%
43 / 147MiMo v2 Flash · 309.8B
25.9%
44 / 148Granite 4.1 30B · 28.9B
25.8%
45 / 150Qwen3 8B · 8.2B
22.6%
46 / 151Gemma 3 27B IT · 27.4B
21.2%
47 / 152Gemma 3 12B IT · 12.2B
17.4%
48 / 153Llama 4 Scout 17B 16E Instruct · 108.6B
17.0%
49 / 154Llama 3.1 8B Instruct · 8.0B
13.2%
50 / 155Phi 4 Mini Instruct · 3.8B
10.8%

Score vs model size

Which models give the most quality for their size — the ones worth running locally.

10B100B1Tmodel size (log scale) →58.7%10.8%Kimi K2.6 · 1T · 53.5%GLM 5.2 · 753B · 50.5%MiMo V2.5 Pro · 1T · 50.2%DeepSeek V4 Pro · 1.6T · 50.0%DeepSeek V4 Pro 0813 · 1.7T · 49.2%Kimi K2.5 · 1T · 49.0%Kimi K2.7 Code · 1T · 47.4%Inkling · 952B · 46.1%GLM 5.3 Flash · 321B · 46.1%MiniMax M3 · 427B · 45.4%GLM 4.7 · 358B · 45.1%DeepSeek V4 Flash · 291B · 44.9%GLM 5.1 · 754B · 43.8%Gemma 4 31B IT · 31B · 43.4%MiMo V2.5 · 311B · 43.1%Ring 2.6 1T · 1T · 42.4%Qwen3 235B A22B Thinking 2507 · 235B · 42.4%Step 3.7 Flash · 201B · 40.1%DeepSeek V3.2 Exp · 685B · 38.9%GPT OSS 120B · 117B · 38.9%GLM 4.6 · 357B · 38.4%Qwen3.6 27B · 28B · 37.3%Trinity Large Thinking · 399B · 36.1%DeepSeek v3 0324 · 685B · 35.8%Qwen3.6 35B A3B · 36B · 35.8%DeepSeek R1 · 685B · 35.7%Qwen3.5 122B A10B · 125B · 35.6%DeepSeek v3 · 685B · 35.4%Qwen3 32B · 33B · 35.4%Qwen3 30B A3B Thinking 2507 · 31B · 33.3%Llama 4 Maverick 17B 128E Instruct · 402B · 33.1%Qwen3 Coder Next · 80B · 32.3%Llama 3.3 70B Instruct · 71B · 26.0%MiMo v2 Flash · 310B · 25.9%Granite 4.1 30B · 29B · 25.8%Gemma 3 27B IT · 27B · 21.2%Gemma 3 12B IT · 12B · 17.4%Llama 4 Scout 17B 16E Instruct · 109B · 17.0%Phi 4 Mini Instruct · 4B · 10.8%Phi 4 Mini InstructLlama 3.1 8B Instruct · 8B · 13.2%Llama 3.1 8B InstructQwen3 8B · 8B · 22.6%Qwen3 8BQwen3.5 9B · 10B · 27.6%Qwen3.5 9BQwen3 14B · 15B · 31.6%Qwen3 14BGPT OSS 20B · 21B · 34.4%GPT OSS 20BQwen3.8 27B · 28B · 44.7%Qwen3.8 27BMiniMax M2.7 · 229B · 47.0%Inkling Small · 266B · 48.7%Inkling SmallDeepSeek V4 Flash 0731 · 304B · 49.9%DeepSeek V4 Flash 0731GLM 5.3 · 753B · 56.5%GLM 5.3Kimi K3 · 2.8T · 58.7%Kimi K3
Each dot is a model. Up = higher score, left = smaller (easier to run locally). The dashed line marks the efficiency frontier — the best score you can get at each size or smaller.
  • Phi 4 Mini Instruct, 4B, score 10.8% — on the efficiency frontier (best score at its size or smaller).
  • Llama 3.1 8B Instruct, 8B, score 13.2% — on the efficiency frontier (best score at its size or smaller).
  • Qwen3 8B, 8B, score 22.6% — on the efficiency frontier (best score at its size or smaller).
  • Qwen3.5 9B, 10B, score 27.6% — on the efficiency frontier (best score at its size or smaller).
  • Qwen3 14B, 15B, score 31.6% — on the efficiency frontier (best score at its size or smaller).
  • GPT OSS 20B, 21B, score 34.4% — on the efficiency frontier (best score at its size or smaller).
  • Qwen3.8 27B, 28B, score 44.7% — on the efficiency frontier (best score at its size or smaller).
  • MiniMax M2.7, 229B, score 47.0% — on the efficiency frontier (best score at its size or smaller).
  • Inkling Small, 266B, score 48.7% — on the efficiency frontier (best score at its size or smaller).
  • DeepSeek V4 Flash 0731, 304B, score 49.9% — on the efficiency frontier (best score at its size or smaller).
  • GLM 5.3, 753B, score 56.5% — on the efficiency frontier (best score at its size or smaller).
  • Kimi K3, 2.8T, score 58.7% — on the efficiency frontier (best score at its size or smaller).

SciCode: frequently asked questions

What is the best open LLM on SciCode?
Kimi K3 is the top open model on SciCode, scoring 58.7%. Among all models tested — including proprietary ones — it ranks #5. The top model overall is Claude Fable 5.1 Max (Anthropic) at 62.0%.
What's the best SciCode model you can run on a 24 GB GPU?
Qwen3.8 27B is the highest-scoring open model that fits in 24 GB at 4-bit quantization (about 15 GB), scoring 44.7% on SciCode.
What's the best SciCode model you can run on a 12 GB GPU?
GPT OSS 20B is the highest-scoring open model that fits in 12 GB at 4-bit quantization (about 12 GB), scoring 34.4% on SciCode.
Can open models match proprietary models on SciCode?
Not quite on SciCode: the strongest proprietary model (Claude Fable 5.1 Max) scores 62.0%, ahead of the best open model (Kimi K3) at 58.7% — but you can run the open one yourself.

Scores aggregated from epoch. llmrun does not run this benchmark — see the source for methodology, or the about benchmarks for what it measures.