Coding

LMArena WebDev Leaderboard

Code Arena (formerly WebDev Arena) is Arena's head-to-head coding leaderboard: two models each build a real web app or UI from the same prompt, and people vote blind on which result is better. The rating shown is a relative Bradley-Terry score, not a percentage — compare it only to other models on this board.

Source: lmarena48 open models ranked+75 proprietaryData through Sep 2026

All models ranked on LMArena WebDev

Proprietary / closed models are shown dimmed — you can't run them locally, but they show where the open field stands.

#ModelScore
1GPT 6 Astra Max · proprietary
1793.2
2Claude Fable 5.1 Max · proprietary
1754.9
3Claude Opus 5 Max · proprietary
1691.1
4Qwen3.8 Max · proprietary
1670.9
5Qwen3.8 Max (Sep 02) · proprietary
1662.2
6Claude Opus 5 High · proprietary
1661.3
7Kimi K3 · 2779.9B
1657.9
8Muse Spark 1.3 Max · proprietary
1656.7
9Qwen3.8 Flash Next · 180.0B
1635.6
10Grok 4.7 Xhigh · proprietary
1632.5
11Hy4 Preview · 780.0B
1627.4
12Claude Fable 5 High · proprietary
1626.8
13Muse Spark 1.3 (xHigh) · proprietary
1623.7
14GLM 5.3 · 753.3B
1619.7
15GPT 5.6 Sol Xhigh (codex Harness) · proprietary
1617.3
16Grok 4.6 High · proprietary
1616.2
17DeepSeek V4.1 Flash · 763.2B
1616.1
18GLM 5.3 Flash · 321.3B
1612.2
19GLM 5.2 · 753.3B
1598.2
20Gemini 3.7 Flash High · proprietary
1595.0
21Qwen3.8 27B · 27.8B
1591.2
22Gemini 3.8 Flash High · proprietary
1583.4
23DeepSeek V4 Pro 0813 · 1650.5B
1580.2
24DeepSeek V4 Flash · 290.9B
1579.9
25Claude Opus 4.7 · proprietary
1557.5
26Claude Opus 4.8 High · proprietary
1556.1
27Claude Opus 4.7 High · proprietary
1555.6
28Grok 4.5 · proprietary
1552.4
29Claude Opus 4.6 High · proprietary
1546.6
30Muse Spark 1.1 · proprietary
1541.9
31Claude Sonnet 5 High · proprietary
1537.8
32Claude Opus 4.8 · proprietary
1537.2
33Claude Opus 4.6 · proprietary
1537.1
34Gemini 3.6 Flash High · proprietary
1536.8
35Muse Spark 1.2 (xHigh) · proprietary
1534.4
36Claude Sonnet 4.6 · proprietary
1520.5
37GPT 5.6 Luna Xhigh (codex Harness) · proprietary
1519.8
38Seed 2.1 Pro Preview · proprietary
1519.2
39GPT 5.6 Terra Xhigh (codex Harness) · proprietary
1518.1
40Qwen3.7 Max (May 17, 2026) · proprietary
1517.1
41GPT 5.5 Xhigh (codex Harness) · proprietary
1510.2
42Hy3 · 298.8B
1508.7
43Kimi K2.6 · 1026.9B
1508.6
44GLM 5.1 · 753.9B
1508.1
45Gemini 3.5 Flash High · proprietary
1500.0
46Claude Opus 4.5 High 32K (Nov 01, 2025) · proprietary
1494.9
47Gemini 3.5 Flash Medium · proprietary
1491.0
48GPT 5.5 High (codex Harness) · proprietary
1486.6
49MiniMax M3 · 427.0B
1485.0
50Qwen3.6 Max Preview · proprietary
1478.8
51MiMo V2.5 Pro · 1023.2B
1476.8
52Kimi K2.7 Code · 1026.9B
1472.3
53Claude Opus 4.5 (Nov 01, 2025) · proprietary
1468.3
54DeepSeek V4 Pro · 1598.8B
1463.5
55GPT 5.4 High (codex Harness) · proprietary
1462.5
56Qwen3.6 Plus · proprietary
1460.5
57GPT 5.5 (codex Harness) · proprietary
1457.7
58Gemini 3.5 Flash Lite · proprietary
1446.9
59Gemini 3.1 Pro Preview · proprietary
1446.7
60GPT 5.4 Medium (codex Harness) · proprietary
1442.7
61Gemini 3 Pro · proprietary
1438.9
62Gemini 3 Flash · proprietary
1438.2
63MiMo V2.5 · 310.8B
1437.4
64Kimi K2.5 · 1026.9B
1436.3
65GLM 5 · 753.9B
1435.8
66GLM 4.7 · 358.3B
1434.9
67Mimo v2 Pro · proprietary
1433.7
68GPT 5 Medium · proprietary
1419.7
69GPT 5.2 · proprietary
1416.8
70Inkling · 952.4B
1411.1
71GPT 5.3 Codex (codex Harness) · proprietary
1409.0
72Inkling Small · 266.0B
1406.7
73GLM 5v Turbo · proprietary
1400.6
74Qwen3.5 397B A17B · 403.4B
1398.8
75MiniMax M2.7 · 228.7B
1397.5
76GPT 5.4 Mini High · proprietary
1397.2
77Claude Sonnet 4.5 High 32K (Sep 29, 2025) · proprietary
1392.8
78GPT 5.1 Medium · proprietary
1392.1
79GPT 5.4 · proprietary
1391.8
80Claude Opus 4.1 (Aug 05, 2025) · proprietary
1389.3
81MiniMax M2.1 · 228.7B
1387.4
82Claude Sonnet 4.5 (Sep 29, 2025) · proprietary
1385.4
83MiniMax M2.5 · 228.7B
1384.0
84Gemini 3 Flash (thinking Minimal) · proprietary
1383.1
85Grok 4.20 Beta 0309 Reasoning · proprietary
1373.4
86Solar Pro4 · proprietary
1364.7
87Gemma 4 31B IT · 31.3B
1364.2
88Gemma 4 26B A4B IT · 25.8B
1361.9
89DeepSeek V3.2 · 685.4B
1360.8
90Muse Glimmer 30B · 29.8B
1359.8
91Qwen3.5 122B A10B · 125.1B
1358.2
92Qwen3.5 27B · 27.8B
1357.8
93Hy3 Preview · 298.8B
1357.0
94Grok 4.3 · proprietary
1357.0
95Laguna M.1 · 225.8B
1347.2
96GPT 5.1 · proprietary
1341.0
97GLM 4.6 · 356.8B
1340.8
98GPT 5.2 Codex · proprietary
1338.7
99GPT 5.1 Codex · proprietary
1336.2
100MiMo v2 Flash · 309.8B
1330.7
101Claude Haiku 4.5 (Oct 01, 2025) · proprietary
1329.5
102Kimi K2 Thinking · 1026.4B
1323.0
103Laguna XS.2 · 33.4B
1302.3
104MiniMax M2 · 228.7B
1298.0
105Qwen3 Coder 480B A35B Instruct · 480.2B
1273.9
106DeepSeek V3.2 Exp · 685.4B
1271.8
107Mistral Medium 3.5 128B · 127.7B
1264.6
108KAT Coder Pro V1 · proprietary
1255.2
109Gemini 3.1 Flash Lite Preview · proprietary
1254.5
110Qwen3.5 35B A3B · 36.0B
1250.3
111GPT 5.1 Codex Mini · proprietary
1244.1
112Grok 4.1 Fast Reasoning · proprietary
1240.7
113Qwen3.5 Flash · proprietary
1239.0
114Trinity Large Thinking · 398.6B
1236.8
115Mistral Large 3 675B Instruct 2512 · 675B
1228.8
116Gemini 2.5 Pro · proprietary
1226.0
117Grok 4.1 Thinking · proprietary
1211.2
118Devstral 2 123B Instruct 2512 · 125.0B
1194.2
119Granite 4.1 8B · 8.8B
1191.6
120Mercury 2 · proprietary
1167.0
121Grok Code Fast 1 · proprietary
1165.3
122Grok 4 Fast Reasoning · proprietary
1161.9
123Devstral Medium 2507 · proprietary
1081.4

Score vs model size

Which models give the most quality for their size — the ones worth running locally.

10B100B1Tmodel size (log scale) →1657.91191.6Hy4 Preview · 780B · 1627.4GLM 5.3 · 753B · 1619.7DeepSeek V4.1 Flash · 763B · 1616.1GLM 5.3 Flash · 321B · 1612.2GLM 5.2 · 753B · 1598.2DeepSeek V4 Pro 0813 · 1.7T · 1580.2DeepSeek V4 Flash · 291B · 1579.9Hy3 · 299B · 1508.7Kimi K2.6 · 1T · 1508.6GLM 5.1 · 754B · 1508.1MiniMax M3 · 427B · 1485.0MiMo V2.5 Pro · 1T · 1476.8Kimi K2.7 Code · 1T · 1472.3DeepSeek V4 Pro · 1.6T · 1463.5MiMo V2.5 · 311B · 1437.4Kimi K2.5 · 1T · 1436.3GLM 5 · 754B · 1435.8GLM 4.7 · 358B · 1434.9Inkling · 952B · 1411.1Inkling Small · 266B · 1406.7Qwen3.5 397B A17B · 403B · 1398.8MiniMax M2.7 · 229B · 1397.5MiniMax M2.1 · 229B · 1387.4MiniMax M2.5 · 229B · 1384.0Gemma 4 31B IT · 31B · 1364.2DeepSeek V3.2 · 685B · 1360.8Muse Glimmer 30B · 30B · 1359.8Qwen3.5 122B A10B · 125B · 1358.2Qwen3.5 27B · 28B · 1357.8Hy3 Preview · 299B · 1357.0Laguna M.1 · 226B · 1347.2GLM 4.6 · 357B · 1340.8MiMo v2 Flash · 310B · 1330.7Kimi K2 Thinking · 1T · 1323.0Laguna XS.2 · 33B · 1302.3MiniMax M2 · 229B · 1298.0Qwen3 Coder 480B A35B Instruct · 480B · 1273.9DeepSeek V3.2 Exp · 685B · 1271.8Mistral Medium 3.5 128B · 128B · 1264.6Qwen3.5 35B A3B · 36B · 1250.3Trinity Large Thinking · 399B · 1236.8Mistral Large 3 675B Instruct 2512 · 675B · 1228.8Devstral 2 123B Instruct 2512 · 125B · 1194.2Granite 4.1 8B · 9B · 1191.6Granite 4.1 8BGemma 4 26B A4B IT · 26B · 1361.9Gemma 4 26B A4B ITQwen3.8 27B · 28B · 1591.2Qwen3.8 27BQwen3.8 Flash Next · 180B · 1635.6Qwen3.8 Flash NextKimi K3 · 2.8T · 1657.9Kimi K3
Each dot is a model. Up = higher score, left = smaller (easier to run locally). The dashed line marks the efficiency frontier — the best score you can get at each size or smaller.
  • Granite 4.1 8B, 9B, score 1191.6 — on the efficiency frontier (best score at its size or smaller).
  • Gemma 4 26B A4B IT, 26B, score 1361.9 — on the efficiency frontier (best score at its size or smaller).
  • Qwen3.8 27B, 28B, score 1591.2 — on the efficiency frontier (best score at its size or smaller).
  • Qwen3.8 Flash Next, 180B, score 1635.6 — on the efficiency frontier (best score at its size or smaller).
  • Kimi K3, 2.8T, score 1657.9 — on the efficiency frontier (best score at its size or smaller).

LMArena WebDev: frequently asked questions

What is the best open LLM on LMArena WebDev?
Kimi K3 is the top open model on LMArena WebDev, scoring 1657.9. Among all models tested — including proprietary ones — it ranks #7. The top model overall is GPT 6 Astra Max (OpenAI) at 1793.2.
What's the best LMArena WebDev model you can run on a 24 GB GPU?
Qwen3.8 27B is the highest-scoring open model that fits in 24 GB at 4-bit quantization (about 15 GB), scoring 1591.2 on LMArena WebDev.
What's the best LMArena WebDev model you can run on a 12 GB GPU?
Granite 4.1 8B is the highest-scoring open model that fits in 12 GB at 4-bit quantization (about 5 GB), scoring 1191.6 on LMArena WebDev.
Can open models match proprietary models on LMArena WebDev?
Not quite on LMArena WebDev: the strongest proprietary model (GPT 6 Astra Max) scores 1793.2, ahead of the best open model (Kimi K3) at 1657.9 — but you can run the open one yourself.

Scores aggregated from lmarena. llmrun does not run this benchmark — see the source for methodology, or the about benchmarks for what it measures.