Coding
LMArena WebDev Leaderboard
Code Arena (formerly WebDev Arena) is Arena's head-to-head coding leaderboard: two models each build a real web app or UI from the same prompt, and people vote blind on which result is better. The rating shown is a relative Bradley-Terry score, not a percentage — compare it only to other models on this board.
Source: lmarena48 open models ranked+75 proprietaryData through Sep 2026
Open models ranked on LMArena WebDev
# shows rank among open models / rank overall (including proprietary).
| # | Model | Score |
|---|---|---|
| 1 / 7 | Kimi K3 · 2779.9B | 1657.9 |
| 2 / 9 | Qwen3.8 Flash Next · 180.0B | 1635.6 |
| 3 / 11 | Hy4 Preview · 780.0B | 1627.4 |
| 4 / 14 | GLM 5.3 · 753.3B | 1619.7 |
| 5 / 17 | DeepSeek V4.1 Flash · 763.2B | 1616.1 |
| 6 / 18 | GLM 5.3 Flash · 321.3B | 1612.2 |
| 7 / 19 | GLM 5.2 · 753.3B | 1598.2 |
| 8 / 21 | Qwen3.8 27B · 27.8B | 1591.2 |
| 9 / 23 | DeepSeek V4 Pro 0813 · 1650.5B | 1580.2 |
| 10 / 24 | DeepSeek V4 Flash · 290.9B | 1579.9 |
| 11 / 42 | Hy3 · 298.8B | 1508.7 |
| 12 / 43 | Kimi K2.6 · 1026.9B | 1508.6 |
| 13 / 44 | GLM 5.1 · 753.9B | 1508.1 |
| 14 / 49 | MiniMax M3 · 427.0B | 1485.0 |
| 15 / 51 | MiMo V2.5 Pro · 1023.2B | 1476.8 |
| 16 / 52 | Kimi K2.7 Code · 1026.9B | 1472.3 |
| 17 / 54 | DeepSeek V4 Pro · 1598.8B | 1463.5 |
| 18 / 63 | MiMo V2.5 · 310.8B | 1437.4 |
| 19 / 64 | Kimi K2.5 · 1026.9B | 1436.3 |
| 20 / 65 | GLM 5 · 753.9B | 1435.8 |
| 21 / 66 | GLM 4.7 · 358.3B | 1434.9 |
| 22 / 70 | Inkling · 952.4B | 1411.1 |
| 23 / 72 | Inkling Small · 266.0B | 1406.7 |
| 24 / 74 | Qwen3.5 397B A17B · 403.4B | 1398.8 |
| 25 / 75 | MiniMax M2.7 · 228.7B | 1397.5 |
| 26 / 81 | MiniMax M2.1 · 228.7B | 1387.4 |
| 27 / 83 | MiniMax M2.5 · 228.7B | 1384.0 |
| 28 / 87 | Gemma 4 31B IT · 31.3B | 1364.2 |
| 29 / 88 | Gemma 4 26B A4B IT · 25.8B | 1361.9 |
| 30 / 89 | DeepSeek V3.2 · 685.4B | 1360.8 |
| 31 / 90 | Muse Glimmer 30B · 29.8B | 1359.8 |
| 32 / 91 | Qwen3.5 122B A10B · 125.1B | 1358.2 |
| 33 / 92 | Qwen3.5 27B · 27.8B | 1357.8 |
| 34 / 93 | Hy3 Preview · 298.8B | 1357.0 |
| 35 / 95 | Laguna M.1 · 225.8B | 1347.2 |
| 36 / 97 | GLM 4.6 · 356.8B | 1340.8 |
| 37 / 100 | MiMo v2 Flash · 309.8B | 1330.7 |
| 38 / 102 | Kimi K2 Thinking · 1026.4B | 1323.0 |
| 39 / 103 | Laguna XS.2 · 33.4B | 1302.3 |
| 40 / 104 | MiniMax M2 · 228.7B | 1298.0 |
| 41 / 105 | Qwen3 Coder 480B A35B Instruct · 480.2B | 1273.9 |
| 42 / 106 | DeepSeek V3.2 Exp · 685.4B | 1271.8 |
| 43 / 107 | Mistral Medium 3.5 128B · 127.7B | 1264.6 |
| 44 / 110 | Qwen3.5 35B A3B · 36.0B | 1250.3 |
| 45 / 114 | Trinity Large Thinking · 398.6B | 1236.8 |
| 46 / 115 | Mistral Large 3 675B Instruct 2512 · 675B | 1228.8 |
| 47 / 118 | Devstral 2 123B Instruct 2512 · 125.0B | 1194.2 |
| 48 / 119 | Granite 4.1 8B · 8.8B | 1191.6 |
Score vs model size
Which models give the most quality for their size — the ones worth running locally.
- Granite 4.1 8B, 9B, score 1191.6 — on the efficiency frontier (best score at its size or smaller).
- Gemma 4 26B A4B IT, 26B, score 1361.9 — on the efficiency frontier (best score at its size or smaller).
- Qwen3.8 27B, 28B, score 1591.2 — on the efficiency frontier (best score at its size or smaller).
- Qwen3.8 Flash Next, 180B, score 1635.6 — on the efficiency frontier (best score at its size or smaller).
- Kimi K3, 2.8T, score 1657.9 — on the efficiency frontier (best score at its size or smaller).
LMArena WebDev: frequently asked questions
- What is the best open LLM on LMArena WebDev?
- Kimi K3 is the top open model on LMArena WebDev, scoring 1657.9. Among all models tested — including proprietary ones — it ranks #7. The top model overall is GPT 6 Astra Max (OpenAI) at 1793.2.
- What's the best LMArena WebDev model you can run on a 24 GB GPU?
- Qwen3.8 27B is the highest-scoring open model that fits in 24 GB at 4-bit quantization (about 15 GB), scoring 1591.2 on LMArena WebDev.
- What's the best LMArena WebDev model you can run on a 12 GB GPU?
- Granite 4.1 8B is the highest-scoring open model that fits in 12 GB at 4-bit quantization (about 5 GB), scoring 1191.6 on LMArena WebDev.
- Can open models match proprietary models on LMArena WebDev?
- Not quite on LMArena WebDev: the strongest proprietary model (GPT 6 Astra Max) scores 1793.2, ahead of the best open model (Kimi K3) at 1657.9 — but you can run the open one yourself.
Scores aggregated from lmarena. llmrun does not run this benchmark — see the source for methodology, or the about benchmarks for what it measures.