Coding

LMArena WebDev Leaderboard

Code Arena (formerly WebDev Arena) is Arena's head-to-head coding leaderboard: two models each build a real web app or UI from the same prompt, and people vote blind on which result is better. The rating shown is a relative Bradley-Terry score, not a percentage — compare it only to other models on this board.

Source: lmarena48 open models ranked+75 proprietaryData through Sep 2026

Open models ranked on LMArena WebDev

# shows rank among open models / rank overall (including proprietary).

#ModelScore
1 / 7Kimi K3 · 2779.9B
1657.9
2 / 9Qwen3.8 Flash Next · 180.0B
1635.6
3 / 11Hy4 Preview · 780.0B
1627.4
4 / 14GLM 5.3 · 753.3B
1619.7
5 / 17DeepSeek V4.1 Flash · 763.2B
1616.1
6 / 18GLM 5.3 Flash · 321.3B
1612.2
7 / 19GLM 5.2 · 753.3B
1598.2
8 / 21Qwen3.8 27B · 27.8B
1591.2
9 / 23DeepSeek V4 Pro 0813 · 1650.5B
1580.2
10 / 24DeepSeek V4 Flash · 290.9B
1579.9
11 / 42Hy3 · 298.8B
1508.7
12 / 43Kimi K2.6 · 1026.9B
1508.6
13 / 44GLM 5.1 · 753.9B
1508.1
14 / 49MiniMax M3 · 427.0B
1485.0
15 / 51MiMo V2.5 Pro · 1023.2B
1476.8
16 / 52Kimi K2.7 Code · 1026.9B
1472.3
17 / 54DeepSeek V4 Pro · 1598.8B
1463.5
18 / 63MiMo V2.5 · 310.8B
1437.4
19 / 64Kimi K2.5 · 1026.9B
1436.3
20 / 65GLM 5 · 753.9B
1435.8
21 / 66GLM 4.7 · 358.3B
1434.9
22 / 70Inkling · 952.4B
1411.1
23 / 72Inkling Small · 266.0B
1406.7
24 / 74Qwen3.5 397B A17B · 403.4B
1398.8
25 / 75MiniMax M2.7 · 228.7B
1397.5
26 / 81MiniMax M2.1 · 228.7B
1387.4
27 / 83MiniMax M2.5 · 228.7B
1384.0
28 / 87Gemma 4 31B IT · 31.3B
1364.2
29 / 88Gemma 4 26B A4B IT · 25.8B
1361.9
30 / 89DeepSeek V3.2 · 685.4B
1360.8
31 / 90Muse Glimmer 30B · 29.8B
1359.8
32 / 91Qwen3.5 122B A10B · 125.1B
1358.2
33 / 92Qwen3.5 27B · 27.8B
1357.8
34 / 93Hy3 Preview · 298.8B
1357.0
35 / 95Laguna M.1 · 225.8B
1347.2
36 / 97GLM 4.6 · 356.8B
1340.8
37 / 100MiMo v2 Flash · 309.8B
1330.7
38 / 102Kimi K2 Thinking · 1026.4B
1323.0
39 / 103Laguna XS.2 · 33.4B
1302.3
40 / 104MiniMax M2 · 228.7B
1298.0
41 / 105Qwen3 Coder 480B A35B Instruct · 480.2B
1273.9
42 / 106DeepSeek V3.2 Exp · 685.4B
1271.8
43 / 107Mistral Medium 3.5 128B · 127.7B
1264.6
44 / 110Qwen3.5 35B A3B · 36.0B
1250.3
45 / 114Trinity Large Thinking · 398.6B
1236.8
46 / 115Mistral Large 3 675B Instruct 2512 · 675B
1228.8
47 / 118Devstral 2 123B Instruct 2512 · 125.0B
1194.2
48 / 119Granite 4.1 8B · 8.8B
1191.6

Score vs model size

Which models give the most quality for their size — the ones worth running locally.

10B100B1Tmodel size (log scale) →1657.91191.6Hy4 Preview · 780B · 1627.4GLM 5.3 · 753B · 1619.7DeepSeek V4.1 Flash · 763B · 1616.1GLM 5.3 Flash · 321B · 1612.2GLM 5.2 · 753B · 1598.2DeepSeek V4 Pro 0813 · 1.7T · 1580.2DeepSeek V4 Flash · 291B · 1579.9Hy3 · 299B · 1508.7Kimi K2.6 · 1T · 1508.6GLM 5.1 · 754B · 1508.1MiniMax M3 · 427B · 1485.0MiMo V2.5 Pro · 1T · 1476.8Kimi K2.7 Code · 1T · 1472.3DeepSeek V4 Pro · 1.6T · 1463.5MiMo V2.5 · 311B · 1437.4Kimi K2.5 · 1T · 1436.3GLM 5 · 754B · 1435.8GLM 4.7 · 358B · 1434.9Inkling · 952B · 1411.1Inkling Small · 266B · 1406.7Qwen3.5 397B A17B · 403B · 1398.8MiniMax M2.7 · 229B · 1397.5MiniMax M2.1 · 229B · 1387.4MiniMax M2.5 · 229B · 1384.0Gemma 4 31B IT · 31B · 1364.2DeepSeek V3.2 · 685B · 1360.8Muse Glimmer 30B · 30B · 1359.8Qwen3.5 122B A10B · 125B · 1358.2Qwen3.5 27B · 28B · 1357.8Hy3 Preview · 299B · 1357.0Laguna M.1 · 226B · 1347.2GLM 4.6 · 357B · 1340.8MiMo v2 Flash · 310B · 1330.7Kimi K2 Thinking · 1T · 1323.0Laguna XS.2 · 33B · 1302.3MiniMax M2 · 229B · 1298.0Qwen3 Coder 480B A35B Instruct · 480B · 1273.9DeepSeek V3.2 Exp · 685B · 1271.8Mistral Medium 3.5 128B · 128B · 1264.6Qwen3.5 35B A3B · 36B · 1250.3Trinity Large Thinking · 399B · 1236.8Mistral Large 3 675B Instruct 2512 · 675B · 1228.8Devstral 2 123B Instruct 2512 · 125B · 1194.2Granite 4.1 8B · 9B · 1191.6Granite 4.1 8BGemma 4 26B A4B IT · 26B · 1361.9Gemma 4 26B A4B ITQwen3.8 27B · 28B · 1591.2Qwen3.8 27BQwen3.8 Flash Next · 180B · 1635.6Qwen3.8 Flash NextKimi K3 · 2.8T · 1657.9Kimi K3
Each dot is a model. Up = higher score, left = smaller (easier to run locally). The dashed line marks the efficiency frontier — the best score you can get at each size or smaller.
  • Granite 4.1 8B, 9B, score 1191.6 — on the efficiency frontier (best score at its size or smaller).
  • Gemma 4 26B A4B IT, 26B, score 1361.9 — on the efficiency frontier (best score at its size or smaller).
  • Qwen3.8 27B, 28B, score 1591.2 — on the efficiency frontier (best score at its size or smaller).
  • Qwen3.8 Flash Next, 180B, score 1635.6 — on the efficiency frontier (best score at its size or smaller).
  • Kimi K3, 2.8T, score 1657.9 — on the efficiency frontier (best score at its size or smaller).

LMArena WebDev: frequently asked questions

What is the best open LLM on LMArena WebDev?
Kimi K3 is the top open model on LMArena WebDev, scoring 1657.9. Among all models tested — including proprietary ones — it ranks #7. The top model overall is GPT 6 Astra Max (OpenAI) at 1793.2.
What's the best LMArena WebDev model you can run on a 24 GB GPU?
Qwen3.8 27B is the highest-scoring open model that fits in 24 GB at 4-bit quantization (about 15 GB), scoring 1591.2 on LMArena WebDev.
What's the best LMArena WebDev model you can run on a 12 GB GPU?
Granite 4.1 8B is the highest-scoring open model that fits in 12 GB at 4-bit quantization (about 5 GB), scoring 1191.6 on LMArena WebDev.
Can open models match proprietary models on LMArena WebDev?
Not quite on LMArena WebDev: the strongest proprietary model (GPT 6 Astra Max) scores 1793.2, ahead of the best open model (Kimi K3) at 1657.9 — but you can run the open one yourself.

Scores aggregated from lmarena. llmrun does not run this benchmark — see the source for methodology, or the about benchmarks for what it measures.