Every LLM benchmark, one honest table.
Scores, prices, and speed for 152 models — aggregated from public sources, with the origin and provenance of every number visible. Refreshed automatically, last 5h ago.
Frontier right now
Full leaderboard →| # | Model | Lab | Context | $/1M in · out | Intelligence | Coding | Agentic | Arena Elo |
|---|---|---|---|---|---|---|---|---|
| 1 | Claude Opus 4.7 | Anthropic | 1M | $5 · $25 | 55.0 | 73.6 | 38.6 | 1490 |
| 2 | Claude Fable 5.1 | Anthropic | 1M | $10 · $50 | 53.4 | 81.6 | 57.9 | — |
| 3 | GPT-5.4 | OpenAI | 1.1M | $2.5 · $15 | 53.1 | 71.1 | 44.2 | 1470 |
| 4 | GPT-6 Astra | OpenAI | 1.1M | $10 · $50 | 52.7 | 76.9 | 51.0 | — |
| 5 | Claude Opus 5 | Anthropic | 1M | $5 · $25 | 50.8 | 78.0 | 56.5 | 1505 |
| 6 | Claude Fable 5 | Anthropic | 1M | $10 · $50 | 49.6 | 76.5 | 50.7 | 1493 |
| 7 | GPT-5.6 Sol | OpenAI | 1.1M | $4 · $20 | 47.0 | 77.4 | 50.2 | 1455 |
| 8 | Qwen3.8 Max | Alibaba | 1M | $2 · $6 | 45.4 | 76.2 | 56.0 | 1481 |
| 9 | GLM-5.3open | Zhipu AI | 1M | $1.4 · $4.4 | 44.8 | 74.8 | 53.1 | — |
| 10 | Grok 4.6 | xAI | 500K | $2 · $6 | 44.3 | 76.8 | 53.0 | 1430 |
| 11 | Kimi K3open | Moonshot AI | 1.0M | $3 · $15 | 43.6 | 76.2 | 50.0 | 1482 |
| 12 | GPT-5.6 Terra | OpenAI | 1.1M | $2 · $12 | 42.1 | 76.7 | 43.2 | 1446 |
| 13 | Claude Opus 4.8 | Anthropic | 1M | $5 · $25 | 41.8 | 74.3 | 41.9 | 1462 |
| 14 | GLM-5.3-Flashopen | Zhipu AI | 1M | $0.15 · $0.5 | 41.8 | 71.5 | 50.9 | 1472 |
| 15 | Gemini 3.8 Flash | 1.0M | $0.75 · $3.75 | 40.9 | 76.3 | 40.2 | 1495 |
Intelligence, coding & agentic indices by Artificial Analysis; Arena Elo by LMArena (CC BY 4.0). Best value across model variants shown.
Intelligence vs. price
The frontier you actually pay for — up and to the left is better. Hover a point for details; click it to open the model. Free-tier models excluded.
API-onlyOpen weights
New models
- GLM-5.3-FlashXZhipu AISep 18, 2026
- DeepSeek V4.1 FlashDeepSeekSep 10, 2026
- DeepSeek V4 Flash Vision ExpDeepSeekSep 10, 2026
- DeepSeek V4 FlashDeepSeekSep 10, 2026
- Mercury 2.5InceptionSep 8, 2026
- GPT-6 AstraOpenAISep 4, 2026
- Muse Spark 1.3 ContributorMetaSep 2, 2026
- Muse Spark 1.3MetaSep 2, 2026
How to read the numbers
- independent — measured by an independent evaluator
- crowd — human preference votes (Elo)
- mirror — mirrored via an aggregator API
- vendor — self-reported by the model's vendor
Every score keeps a link to where it came from and when we saw it. Read the methodology.