llm.ing

Every LLM benchmark, one honest table.

Scores, prices, and speed for 152 models — aggregated from public sources, with the origin and provenance of every number visible. Refreshed automatically, last 5h ago.

Frontier right now

Full leaderboard →
#ModelLabContext$/1M in · outIntelligenceCodingAgenticArena Elo
1Claude Opus 4.7Anthropic1M$5 · $2555.073.638.61490
2Claude Fable 5.1Anthropic1M$10 · $5053.481.657.9
3GPT-5.4OpenAI1.1M$2.5 · $1553.171.144.21470
4GPT-6 AstraOpenAI1.1M$10 · $5052.776.951.0
5Claude Opus 5Anthropic1M$5 · $2550.878.056.51505
6Claude Fable 5Anthropic1M$10 · $5049.676.550.71493
7GPT-5.6 SolOpenAI1.1M$4 · $2047.077.450.21455
8Qwen3.8 MaxAlibaba1M$2 · $645.476.256.01481
9GLM-5.3openZhipu AI1M$1.4 · $4.444.874.853.1
10Grok 4.6xAI500K$2 · $644.376.853.01430
11Kimi K3openMoonshot AI1.0M$3 · $1543.676.250.01482
12GPT-5.6 TerraOpenAI1.1M$2 · $1242.176.743.21446
13Claude Opus 4.8Anthropic1M$5 · $2541.874.341.91462
14GLM-5.3-FlashopenZhipu AI1M$0.15 · $0.541.871.550.91472
15Gemini 3.8 FlashGoogle1.0M$0.75 · $3.7540.976.340.21495

Intelligence, coding & agentic indices by Artificial Analysis; Arena Elo by LMArena (CC BY 4.0). Best value across model variants shown.

Intelligence vs. price

The frontier you actually pay for — up and to the left is better. Hover a point for details; click it to open the model. Free-tier models excluded.

API-onlyOpen weights
10152025303540455055↑ AA Intelligence Index$0.1$0.2$0.3$1$2$3$10input price, $ per 1M tokens →DeepSeek R1 $0.5/1M in · index 11.4DeepSeek V3.2 $0.269/1M in · index 32.6DeepSeek V4 Flash $0.15/1M in · index 34.3DeepSeek V4 Pro $0.435/1M in · index 36.0Claude Sonnet 4.6 $3/1M in · index 30.1Claude Haiku 4.5 (latest) $1/1M in · index 16.9Claude Opus 4.6 $5/1M in · index 26.4Claude Fable 5 $10/1M in · index 49.6Claude Opus 4.8 $5/1M in · index 41.8Claude Sonnet 4.5 (latest) $3/1M in · index 20.7Claude Opus 4.7 $5/1M in · index 55.0Claude Opus 4.5 (latest) $5/1M in · index 29.1Claude Sonnet 5 $2/1M in · index 38.2GPT-4.1 mini $0.4/1M in · index 14.8GPT-5.6 Sol $4/1M in · index 47.0GPT-4.1 nano $0.1/1M in · index 9.6o3-mini $1.1/1M in · index 12.5GPT-5.5 $5/1M in · index 38.4GPT-5 $1.25/1M in · index 35.3GPT-5.4 $2.5/1M in · index 53.1GPT-5.2 Codex $1.75/1M in · index 28.5GPT-5.4 nano $0.2/1M in · index 20.7GPT-5.4 mini $0.75/1M in · index 24.1GPT-5.6 Luna $0.2/1M in · index 37.3GPT-5.2 $1.75/1M in · index 30.4GPT-5 Mini $0.25/1M in · index 16.8GPT-5.1 $1.25/1M in · index 37.5GPT-5.1 Codex $1.25/1M in · index 23.7o3 $2/1M in · index 20.2GPT-5.6 Terra $2/1M in · index 42.1GPT-4.1 $2/1M in · index 12.7Gemini 3.5 Flash $1.5/1M in · index 32.6Gemini 2.5 Flash $0.3/1M in · index 13.1Gemini 3.5 Flash Lite $0.3/1M in · index 22.2Gemini 3.1 Flash Lite Preview $0.25/1M in · index 15.6Gemma 4 26B A4B IT $0.09/1M in · index 26.1Gemini 3.6 Flash $0.75/1M in · index 34.0Gemma 4 31B IT $0.09/1M in · index 15.4Gemini 3.1 Pro Preview $2/1M in · index 29.7Gemini 2.5 Pro $1.25/1M in · index 16.1Gemini 2.5 Flash-Lite $0.1/1M in · index 6.7Grok 4.3 $1.25/1M in · index 24.9Grok 4.5 $2/1M in · index 38.8Grok Build 0.1 $1/1M in · index 40.7Muse Spark 1.1 $1.25/1M in · index 33.7Devstral 2 $0.4/1M in · index 8.6Magistral Small $0.5/1M in · index 8.2Qwen3-Coder 480B-A35B Instruct $1.5/1M in · index 11.9Qwen3.7 Plus $0.5/1M in · index 25.2Qwen3 32B $0.7/1M in · index 7.2Qwen3.6 35B-A3B $0.248/1M in · index 18.2Qwen3.5 27B $0.3/1M in · index 30.0Qwen3.7 Max $2.5/1M in · index 29.5Qwen3-Next 80B-A3B (Thinking) $0.5/1M in · index 16.9Qwen3.6 27B $0.6/1M in · index 21.4Qwen3.5 35B-A3B $0.25/1M in · index 24.3Qwen3 14B $0.35/1M in · index 6.4Qwen2.5 32B Instruct $0.7/1M in · index 6.9Qwen3 235B-A22B $0.7/1M in · index 12.7Qwen3 8B $0.18/1M in · index 5.2Qwen Turbo $0.05/1M in · index 6.4Qwen3.5 397B-A17B $0.6/1M in · index 21.4Qwen3.6 Plus $0.5/1M in · index 40.5Qwen3.5 122B-A10B $0.4/1M in · index 17.7Qwen3-Next 80B-A3B Instruct $0.5/1M in · index 9.6Kimi K2.5 $0.6/1M in · index 36.0Kimi K2.6 $0.95/1M in · index 27.0Kimi K2.7 Code $0.95/1M in · index 25.8Kimi K2 Thinking $0.6/1M in · index 17.2Kimi K3 $3/1M in · index 43.6GLM-5 $1/1M in · index 27.9GLM-5.1 $1.4/1M in · index 26.1GLM-5.2 $1.4/1M in · index 33.7GLM-5V-Turbo $5/1M in · index 23.5GLM-4.7 $0.6/1M in · index 34.5GLM-4.5V $0.6/1M in · index 7.6GLM-4.5 $0.6/1M in · index 12.8GLM-4.6 $0.6/1M in · index 29.3GLM-4.6V $0.3/1M in · index 8.4MiniMax-M2.7 $0.3/1M in · index 38.9MiniMax-M2.1 $0.3/1M in · index 20.9MiniMax-M2.5 $0.3/1M in · index 22.8MiniMax-M3 $0.3/1M in · index 29.2Nemotron 3 Super $0.2/1M in · index 12.8Nemotron 3 Ultra 550B A55B $0.5/1M in · index 22.9Sonar $1/1M in · index 8.7Sonar Pro $3/1M in · index 7.6Sonar Reasoning Pro $2/1M in · index 11.8Mercury 2 $0.25/1M in · index 13.8Step 3.7 Flash $0.185/1M in · index 30.9Claude Opus 5 $5/1M in · index 50.8Qwen3.8 Max $2/1M in · index 45.4Muse Spark 1.2 $1.25/1M in · index 39.6Grok 4.6 $2/1M in · index 44.3Gemini 3.7 Flash $0.75/1M in · index 39.1GLM-5.3 $1.4/1M in · index 44.8GLM-5.3-Flash $0.15/1M in · index 41.8Claude Fable 5.1 $10/1M in · index 53.4Gemini 3.8 Flash $0.75/1M in · index 40.9GPT-6 Astra $10/1M in · index 52.7Claude Opus 4.7Claude Fable 5.1GPT-5.4GPT-6 AstraClaude Opus 5Claude Fable 5

New models

How to read the numbers

  • independent — measured by an independent evaluator
  • crowd — human preference votes (Elo)
  • mirror — mirrored via an aggregator API
  • vendor — self-reported by the model's vendor

Every score keeps a link to where it came from and when we saw it. Read the methodology.