llm.ing

LMArena WebDev

crowdrun by LMArenaunit: eloCC-BY-4.0observed Jul 21, 202650 modelssource ↗methodology ↗

Caveat: Head-to-head web app building votes.

TextWebDevVision
#ModelScore
1Kimi K3
1678
±17· 1.8K votes
2Claude Fable 5
1634
±12· 3K votes
3Claude Opus 4.8thinking
1565
±8· 8K votes
4Claude Opus 4.7thinking
1559
±7· 11.2K votes
5Grok 4.5
1557
±12· 2.9K votes
6Claude Sonnet 5high-effort
1544
±11· 3.6K votes
7Claude Opus 4.6thinking
1542
±6· 13.5K votes
8Muse Spark 1.1
1539
±12· 2.7K votes
9Gemini 3.6 Flash
1537
±13· 2.3K votes
10GLM-5.1
1526
±8· 6.5K votes
11Claude Sonnet 4.6
1523
±6· 16.9K votes
12Kimi K2.6
1515
±7· 9.4K votes
13Qwen3.7 Max
1515
±8· 7K votes
14MiniMax-M3
1491
±8· 7K votes
15Claude Opus 4.5 (latest)thinking-32k
1490
±7· 13.1K votes
16Qwen3.6 Max Preview
1480
±12· 2.5K votes
17Kimi K2.7 Code
1470
±9· 4.7K votes
18DeepSeek V4 Prothinking
1459
±7· 9.6K votes
19Qwen3.6 Plus
1458
±6· 12.6K votes
20Gemini 3.1 Pro Preview
1444
±5· 18.2K votes
21GLM-4.7
1440
±10· 4.9K votes
22Kimi K2.5thinking
1433
±6· 15.7K votes
23GLM-5
1430
±8· 7.5K votes
24GPT-5.2
1406
±17· 1.5K votes
25GLM-5V-Turbo
1401
±16· 1.6K votes
26MiniMax-M2.7
1397
±6· 11.2K votes
27GPT-5.4 minihigh-effort
1397
±7· 10.7K votes
28Qwen3.5 397B-A17B
1396
±6· 15.3K votes
29Claude Sonnet 4.5 (latest)thinking-32k
1388
±7· 15.7K votes
30GPT-5.4
1386
±19· 1K votes
31Claude Opus 4.1 (latest)
1386
±9· 8.6K votes
32MiniMax-M2.5
1381
±8· 7.9K votes
33DeepSeek V3.2thinking
1368
±8· 7.9K votes
34Qwen3.5 122B-A10B
1364
±7· 8.2K votes
35Grok 4.3
1360
±7· 9K votes
36Qwen3.5 27B
1357
±8· 7.7K votes
37GLM-4.6
1356
±9· 8.3K votes
38GPT-5.1
1339
±7· 12.9K votes
39GPT-5.2 Codex
1334
±8· 7.8K votes
40GPT-5.1 Codex
1330
±10· 6.2K votes
41Kimi K2 Thinking Turbo
1330
±6· 15.4K votes
42Claude Haiku 4.5 (latest)
1327
±5· 26.5K votes
43MiniMax-M2
1305
±9· 8.4K votes
44Qwen3-Coder 480B-A35B Instruct
1281
±7· 15.2K votes
45Gemini 3.1 Flash Lite Preview
1253
±7· 13.6K votes
46Qwen3.5 35B-A3B
1250
±16· 1.8K votes
47GPT-5.1 Codex mini
1240
±18· 1.4K votes
48Gemini 2.5 Pro
1204
±13· 3.3K votes
49Mercury 2
1164
±23· 947 votes
50Devstral Medium
1093
±23· 993 votes

Dot = elo on a 1038–1727 scale (position, not length); whisker = rating confidence interval. Scores link to model pages; every number keeps its source on the model page.