LMArena WebDev
Caveat: Head-to-head web app building votes.
| # | Model | Score | |
|---|---|---|---|
| 1 | Kimi K3Moonshot AI | 1678 ±17· 1.8K votes | |
| 2 | Claude Fable 5Anthropic | 1634 ±12· 3K votes | |
| 3 | Claude Opus 4.8thinkingAnthropic | 1565 ±8· 8K votes | |
| 4 | Claude Opus 4.7thinkingAnthropic | 1559 ±7· 11.2K votes | |
| 5 | Grok 4.5xAI | 1557 ±12· 2.9K votes | |
| 6 | Claude Sonnet 5high-effortAnthropic | 1544 ±11· 3.6K votes | |
| 7 | Claude Opus 4.6thinkingAnthropic | 1542 ±6· 13.5K votes | |
| 8 | Muse Spark 1.1Meta | 1539 ±12· 2.7K votes | |
| 9 | Gemini 3.6 FlashGoogle | 1537 ±13· 2.3K votes | |
| 10 | GLM-5.1Zhipu AI | 1526 ±8· 6.5K votes | |
| 11 | Claude Sonnet 4.6Anthropic | 1523 ±6· 16.9K votes | |
| 12 | Kimi K2.6Moonshot AI | 1515 ±7· 9.4K votes | |
| 13 | Qwen3.7 MaxAlibaba | 1515 ±8· 7K votes | |
| 14 | MiniMax-M3MiniMax (minimax.io) | 1491 ±8· 7K votes | |
| 15 | Claude Opus 4.5 (latest)thinking-32kAnthropic | 1490 ±7· 13.1K votes | |
| 16 | Qwen3.6 Max PreviewAlibaba | 1480 ±12· 2.5K votes | |
| 17 | Kimi K2.7 CodeMoonshot AI | 1470 ±9· 4.7K votes | |
| 18 | DeepSeek V4 ProthinkingDeepSeek | 1459 ±7· 9.6K votes | |
| 19 | Qwen3.6 PlusAlibaba | 1458 ±6· 12.6K votes | |
| 20 | Gemini 3.1 Pro PreviewGoogle | 1444 ±5· 18.2K votes | |
| 21 | GLM-4.7Zhipu AI | 1440 ±10· 4.9K votes | |
| 22 | Kimi K2.5thinkingMoonshot AI | 1433 ±6· 15.7K votes | |
| 23 | GLM-5Zhipu AI | 1430 ±8· 7.5K votes | |
| 24 | GPT-5.2OpenAI | 1406 ±17· 1.5K votes | |
| 25 | GLM-5V-TurboZhipu AI | 1401 ±16· 1.6K votes | |
| 26 | MiniMax-M2.7MiniMax (minimax.io) | 1397 ±6· 11.2K votes | |
| 27 | GPT-5.4 minihigh-effortOpenAI | 1397 ±7· 10.7K votes | |
| 28 | Qwen3.5 397B-A17BAlibaba | 1396 ±6· 15.3K votes | |
| 29 | Claude Sonnet 4.5 (latest)thinking-32kAnthropic | 1388 ±7· 15.7K votes | |
| 30 | GPT-5.4OpenAI | 1386 ±19· 1K votes | |
| 31 | Claude Opus 4.1 (latest)Anthropic | 1386 ±9· 8.6K votes | |
| 32 | MiniMax-M2.5MiniMax (minimax.io) | 1381 ±8· 7.9K votes | |
| 33 | DeepSeek V3.2thinkingDeepSeek | 1368 ±8· 7.9K votes | |
| 34 | Qwen3.5 122B-A10BAlibaba | 1364 ±7· 8.2K votes | |
| 35 | Grok 4.3xAI | 1360 ±7· 9K votes | |
| 36 | Qwen3.5 27BAlibaba | 1357 ±8· 7.7K votes | |
| 37 | GLM-4.6Zhipu AI | 1356 ±9· 8.3K votes | |
| 38 | GPT-5.1OpenAI | 1339 ±7· 12.9K votes | |
| 39 | GPT-5.2 CodexOpenAI | 1334 ±8· 7.8K votes | |
| 40 | GPT-5.1 CodexOpenAI | 1330 ±10· 6.2K votes | |
| 41 | Kimi K2 Thinking TurboMoonshot AI | 1330 ±6· 15.4K votes | |
| 42 | Claude Haiku 4.5 (latest)Anthropic | 1327 ±5· 26.5K votes | |
| 43 | MiniMax-M2MiniMax (minimax.io) | 1305 ±9· 8.4K votes | |
| 44 | Qwen3-Coder 480B-A35B InstructAlibaba | 1281 ±7· 15.2K votes | |
| 45 | Gemini 3.1 Flash Lite PreviewGoogle | 1253 ±7· 13.6K votes | |
| 46 | Qwen3.5 35B-A3BAlibaba | 1250 ±16· 1.8K votes | |
| 47 | GPT-5.1 Codex miniOpenAI | 1240 ±18· 1.4K votes | |
| 48 | Gemini 2.5 ProGoogle | 1204 ±13· 3.3K votes | |
| 49 | Mercury 2Inception | 1164 ±23· 947 votes | |
| 50 | Devstral MediumMistral | 1093 ±23· 993 votes |
Dot = elo on a 1038–1727 scale (position, not length); whisker = rating confidence interval. Scores link to model pages; every number keeps its source on the model page.