Best open LLMs for coding

Open models ranked by coding, from SWE-bench, LiveCodeBench, Aider and other coding boards. The score shown is the overall llmrun Score; the highlighted column is the coding skill score used for the order.

40 benchmarks4 sourcesUpdated 3 Oct 2026Reference scale 2026-Q4How the score is built

Models ranked by coding llmrun Score
#Modelllmrun ScoreCodingAgentsMathScienceReasoningBenchmarksVRAM
1Kimi K31674 GB · 26 benchmarks7361815263261674 GB
2GLM 5.3456 GB · 21 benchmarks676872495821456 GB
3DeepSeek V4 Pro 0813991 GB · 21 benchmarks63–71466521991 GB
4DeepSeek V4 Flash 0731183 GB · 21 benchmarks59–71456221183 GB
5GLM 5.2456 GB · 24 benchmarks586068475224456 GB
6Kimi K2.6620 GB · 18 benchmarks57–63435118620 GB
7DeepSeek V4.1 Flash458 GB · 10 benchmarks54–76455710458 GB
8MiMo V2.5 Pro615 GB · 6 benchmarks54––38426615 GB
9Kimi K2.7 Code620 GB · 19 benchmarks534661405019620 GB
10GLM 5.1457 GB · 13 benchmarks535457404613457 GB
11GLM 5.3 Flash195 GB · 17 benchmarks526362443717195 GB
12MiniMax M3257 GB · 19 benchmarks524046413819257 GB
13Qwen3.8 27B17.4 GB · 9 benchmarks51–573750917.4 GB
14DeepSeek V4 Pro960 GB · 22 benchmarks49–57444922960 GB
15Kimi K2.5620 GB · 20 benchmarks4937–404520620 GB
16GLM 5457 GB · 15 benchmarks4945––4015457 GB
17MiniMax M2.5138 GB · 11 benchmarks4737––4511138 GB
18DeepSeek V3.1 Terminus415 GB · 5 benchmarks47––31395415 GB
19Gemma 4 31B IT20.4 GB · 11 benchmarks47––32401120.4 GB
20DeepSeek R1 0528415 GB · 11 benchmarks46–44–3211415 GB
21DeepSeek V3.2 Exp415 GB · 10 benchmarks46––304410415 GB
22Step 3.7 Flash133 GB · 3 benchmarks—46––30–3133 GB
23DeepSeek V4 Flash175 GB · 10 benchmarks46––364310175 GB
24Qwen3.6 35B A3B22.0 GB · 12 benchmarks44–4935381222.0 GB
25Inkling572 GB · 21 benchmarks43–52405321572 GB
26Gemma 4 26B A4B IT16.1 GB · 9 benchmarks43––3033916.1 GB
27Kimi K2 Thinking620 GB · 9 benchmarks4334–––9620 GB
28MiMo V2.5187 GB · 4 benchmarks43––33–4187 GB
29NVIDIA Nemotron 3 Ultra 550B A55B BF16370 GB · 14 benchmarks43–52374614370 GB
30Qwen3 235B A22B Thinking 2507142 GB · 13 benchmarks42––343713142 GB
31Qwen3.6 27B17.4 GB · 12 benchmarks42–5336371217.4 GB
32MiniMax M2.7138 GB · 6 benchmarks41––35–6138 GB
33Qwen3 235B A22B Instruct 2507142 GB · 8 benchmarks41–––318142 GB
34Kimi K2 Instruct620 GB · 10 benchmarks4128––2510620 GB
35GLM 4.7216 GB · 13 benchmarks413251363513216 GB
36Qwen3 235B A22B142 GB · 10 benchmarks41–––3110142 GB
37GPT OSS 20B12.9 GB · 11 benchmarks40––24261112.9 GB
38Ring 2.6 1T677 GB · 3 benchmarks—40––33–3677 GB
39Qwen3 Coder 480B A35B Instruct289 GB · 6 benchmarks—4031–––6289 GB
40MiMo v2 Flash186 GB · 3 benchmarks—40––17–3186 GB
41Kimi K2 Instruct 0905620 GB · 6 benchmarks40––––6620 GB
42DeepSeek R1415 GB · 14 benchmarks39–40282914415 GB
43DeepSeek v3 0324415 GB · 13 benchmarks39–33272413415 GB
44GLM 4.5216 GB · 7 benchmarks3832–––7216 GB
45Qwen3.5 27B17.4 GB · 6 benchmarks38–––41617.4 GB
46GPT OSS 120B70.5 GB · 20 benchmarks3714–31302070.5 GB
47DeepSeek v3415 GB · 10 benchmarks37–26222110415 GB
48GLM 4.6215 GB · 6 benchmarks3631–29–6215 GB
49NVIDIA Nemotron 3 Super 120B A12B BF1674.7 GB · 4 benchmarks36––27–474.7 GB
50DeepSeek V3.2415 GB · 13 benchmarks3536––4313415 GB
51Qwen3 Coder Next48.2 GB · 3 benchmarks—35––23–348.2 GB
52Qwen3 32B20.3 GB · 9 benchmarks33––2626920.3 GB
53Mistral Large 3 675B Instruct 2512446 GB · 5 benchmarks33––26255446 GB
54Trinity Large Thinking263 GB · 3 benchmarks—32––26–3263 GB
55Llama 4 Maverick 17B 128E Instruct265 GB · 17 benchmarks25–29262317265 GB
56Llama 3.3 70B Instruct46.6 GB · 13 benchmarks22–1917211346.6 GB
57Gemma 3 27B IT18.1 GB · 11 benchmarks18–2916171118.1 GB
58Llama 3.1 8B Instruct5.3 GB · 11 benchmarks14–14716115.3 GB

The range after each score is a 90% interval: given the boards a model has results on, its true score is very likely inside it. Short ranges mean many agreeing results; long ranges mean few or conflicting ones.

Why some models are missing. A model gets a score only with results on at least 4 benchmarks across at least 2 categories, one of them reasoning or coding. Skill scores need at least 2 benchmarks in that skill; otherwise they show as –.

VRAM is llmrun's estimate at Q4_K_M (or the smallest quantization we track) for the weights plus a working context. Proprietary models can't run locally, so the VRAM filter hides them.

How the llmrun Score is built · llmrun does not run these benchmarks; scores are aggregated from public sources.