llm.ing

Every LLM benchmark, one honest table.

Scores, prices, and speed for 129 models — aggregated from public sources, with the origin and provenance of every number visible. Refreshed automatically, last 2h ago.

Frontier right now

Full leaderboard →
#ModelLabContext$/1M in · outIntelligenceCodingAgenticArena Elo
1Claude Fable 5Anthropic1M$10 · $5059.976.552.81504
2GPT-5.6 SolOpenAI1.1M$5 · $3058.977.454.01486
3Kimi K3openMoonshot AI1.0M$3 · $1557.176.250.11482
4Claude Opus 4.8Anthropic1M$5 · $2555.774.347.21472
5GPT-5.6 TerraOpenAI1.1M$2.5 · $1555.076.747.4
6GPT-5.5OpenAI1.1M$5 · $3054.874.944.91490
7Grok 4.5xAI500K$2 · $653.872.445.71478
8Claude Opus 4.7Anthropic1M$5 · $2553.573.644.41499
9Claude Sonnet 5Anthropic1M$2 · $1053.471.546.71460
10GPT-5.4OpenAI1.1M$2.5 · $1551.471.141.11499
11GPT-5.6 LunaOpenAI1.1M$1 · $651.271.445.6
12GLM-5.2Zhipu AI1M$1.4 · $4.451.168.843.1
13Muse Spark 1.1Meta1M$1.25 · $4.2550.671.337.51491
14Gemini 3.5 FlashGoogle1.0M$1.5 · $950.270.137.41490
15Gemini 3.6 FlashGoogle1.0M$1.5 · $7.550.169.238.71491

Intelligence, coding & agentic indices by Artificial Analysis; Arena Elo by LMArena (CC BY 4.0). Best value across model variants shown.

Intelligence vs. price

The frontier you actually pay for — up and to the left is better. Hover a point for details; click it to open the model. Free-tier models excluded.

API-onlyOpen weights
10152025303540455055↑ AA Intelligence Index$0.1$0.2$0.3$1$2$3$10input price, $ per 1M tokens →DeepSeek R1 $0.5/1M in · index 18.5DeepSeek V3.2 $0.269/1M in · index 32.0DeepSeek V4 Flash $0.14/1M in · index 40.3DeepSeek V4 Pro $0.435/1M in · index 44.3Claude Sonnet 4.6 $3/1M in · index 47.2Claude Haiku 4.5 (latest) $1/1M in · index 29.6Claude Opus 4.6 $5/1M in · index 37.8Claude Fable 5 $10/1M in · index 59.9Claude Opus 4.8 $5/1M in · index 55.7Claude Sonnet 4.5 (latest) $3/1M in · index 36.4Claude Opus 4.7 $5/1M in · index 53.5Claude Opus 4.5 (latest) $5/1M in · index 40.8Claude Sonnet 5 $2/1M in · index 53.4GPT-4.1 mini $0.4/1M in · index 14.8GPT-5.6 Sol $5/1M in · index 58.9GPT-4.1 nano $0.1/1M in · index 9.6o3-mini $1.1/1M in · index 19.0GPT-5.5 $5/1M in · index 54.8GPT-5 $1.25/1M in · index 34.7GPT-5.4 $2.5/1M in · index 51.4GPT-5.2 Codex $1.75/1M in · index 40.1GPT-5.4 nano $0.2/1M in · index 38.2GPT-5.4 mini $0.75/1M in · index 40.0GPT-5.6 Luna $1/1M in · index 51.2GPT-5.2 $1.75/1M in · index 42.2GPT-5 Mini $0.25/1M in · index 25.3GPT-5.1 $1.25/1M in · index 36.9GPT-5.1 Codex $1.25/1M in · index 34.7o3 $2/1M in · index 30.4GPT-5.6 Terra $2.5/1M in · index 55.0GPT-4.1 $2/1M in · index 19.4Gemini 3.5 Flash $1.5/1M in · index 50.2Gemini 2.5 Flash $0.3/1M in · index 20.1Gemini 3.5 Flash Lite $0.3/1M in · index 36.5Gemini 3.1 Flash Lite Preview $0.25/1M in · index 25.0Gemma 4 26B A4B IT $0.12/1M in · index 25.7Gemini 3.6 Flash $1.5/1M in · index 50.1Gemma 4 31B IT $0.12/1M in · index 29.4Gemini 3.1 Pro Preview $2/1M in · index 46.5Gemini 2.5 Pro $1.25/1M in · index 25.8Gemini 2.5 Flash-Lite $0.1/1M in · index 6.9Grok 4.3 $1.25/1M in · index 37.6Grok 4.5 $2/1M in · index 53.8Grok Build 0.1 $1/1M in · index 39.8Muse Spark 1.1 $1.25/1M in · index 50.6Devstral 2 $0.4/1M in · index 19.2Magistral Small $0.5/1M in · index 10.7Qwen3-Coder 480B-A35B Instruct $1.5/1M in · index 18.0Qwen3.7 Plus $0.5/1M in · index 39.0Qwen3 32B $0.7/1M in · index 11.5Qwen3.6 35B-A3B $0.248/1M in · index 31.6Qwen3.5 27B $0.3/1M in · index 33.8Qwen3.7 Max $2.5/1M in · index 46.0Qwen3-Next 80B-A3B (Thinking) $0.5/1M in · index 16.7Qwen3.6 27B $0.6/1M in · index 37.1Qwen3.5 35B-A3B $0.25/1M in · index 29.3Qwen3 14B $0.35/1M in · index 10.4Qwen2.5 32B Instruct $0.7/1M in · index 7.5Qwen3 235B-A22B $0.7/1M in · index 19.6Qwen3 8B $0.18/1M in · index 8.3Qwen Turbo $0.05/1M in · index 6.3Qwen3.5 397B-A17B $0.6/1M in · index 33.7Qwen3.6 Plus $0.5/1M in · index 39.6Qwen3.5 122B-A10B $0.4/1M in · index 32.3Qwen3-Next 80B-A3B Instruct $0.5/1M in · index 13.7Kimi K2.5 $0.6/1M in · index 35.4Kimi K2.6 $0.95/1M in · index 44.2Kimi K2.7 Code $0.95/1M in · index 41.9Kimi K2 Thinking $0.6/1M in · index 17.3Kimi K3 $3/1M in · index 57.1GLM-5 $1/1M in · index 39.5GLM-5.1 $1.4/1M in · index 40.2GLM-5.2 $1.4/1M in · index 51.1GLM-5V-Turbo $5/1M in · index 34.5GLM-4.7 $0.6/1M in · index 33.7GLM-4.5V $0.6/1M in · index 9.1GLM-4.5 $0.6/1M in · index 19.5GLM-4.6 $0.6/1M in · index 28.7GLM-4.6V $0.3/1M in · index 11.0MiniMax-M2.7 $0.3/1M in · index 38.1MiniMax-M2.1 $0.3/1M in · index 31.4MiniMax-M2.5 $0.3/1M in · index 33.7MiniMax-M3 $0.3/1M in · index 44.4Nemotron 3 Super $0.2/1M in · index 25.4Nemotron 3 Ultra 550B A55B $0.5/1M in · index 37.8Sonar $1/1M in · index 11.7Sonar Pro $3/1M in · index 9.3Sonar Reasoning Pro $2/1M in · index 17.8Mercury 2 $0.25/1M in · index 21.4Step 3.7 Flash $0.185/1M in · index 30.3Claude Fable 5GPT-5.6 SolKimi K3Claude Opus 4.8GPT-5.6 TerraGPT-5.5

New models

How to read the numbers

  • independent — measured by an independent evaluator
  • crowd — human preference votes (Elo)
  • mirror — mirrored via an aggregator API
  • vendor — self-reported by the model's vendor

Every score keeps a link to where it came from and when we saw it. Read the methodology.