llm.ing

Every LLM benchmark, one honest table.

Scores, prices, and speed for 151 models — aggregated from public sources, with the origin and provenance of every number visible. Refreshed automatically, last 6h ago.

Frontier right now

Full leaderboard →
#ModelLabContext$/1M in · outIntelligenceCodingAgenticArena Elo
1Claude Fable 5.1Anthropic1M$10 · $5065.781.661.3
2Claude Opus 5Anthropic1M$5 · $2563.178.059.21504
3Claude Fable 5Anthropic1M$10 · $5062.176.556.61494
4GPT-5.6 SolOpenAI1.1M$4 · $2060.977.457.81455
5Grok 4.6xAI500K$2 · $660.976.858.71444
6Kimi K3openMoonshot AI1.0M$3 · $1559.776.254.31482
7GLM-5.3openZhipu AI1M$1.4 · $4.459.574.859.1
8Gemini 3.8 FlashGoogle1.0M$0.75 · $3.7558.776.350.0
9GLM-5.3-FlashZhipu AI1M$0.075 · $0.2557.571.558.21471
10Claude Opus 4.8Anthropic1M$5 · $2557.374.349.41462
11Muse Spark 1.2Meta1.0M$1.25 · $4.2556.872.249.3
12GPT-5.6 TerraOpenAI1.1M$2 · $1256.676.750.21447
13GPT-5.5OpenAI1.1M$5 · $3056.374.947.41472
14Gemini 3.7 FlashGoogle1.0M$0.75 · $3.7556.076.145.11491
15Grok 4.5xAI500K$2 · $655.872.448.91452

Intelligence, coding & agentic indices by Artificial Analysis; Arena Elo by LMArena (CC BY 4.0). Best value across model variants shown.

Intelligence vs. price

The frontier you actually pay for — up and to the left is better. Hover a point for details; click it to open the model. Free-tier models excluded.

API-onlyOpen weights
101520253035404550556065↑ AA Intelligence Index$0.1$0.2$0.3$1$2$3$10input price, $ per 1M tokens →DeepSeek R1 $0.5/1M in · index 18.6DeepSeek V3.2 $0.269/1M in · index 32.6DeepSeek V4 Flash $0.14/1M in · index 51.8DeepSeek V4 Pro $0.435/1M in · index 53.2Claude Sonnet 4.6 $3/1M in · index 48.4Claude Haiku 4.5 (latest) $1/1M in · index 29.9Claude Opus 4.6 $5/1M in · index 38.8Claude Fable 5 $10/1M in · index 62.1Claude Opus 4.8 $5/1M in · index 57.3Claude Sonnet 4.5 (latest) $3/1M in · index 37.4Claude Opus 4.7 $5/1M in · index 55.0Claude Opus 4.5 (latest) $5/1M in · index 41.9Claude Sonnet 5 $2/1M in · index 55.3GPT-4.1 mini $0.4/1M in · index 14.8GPT-5.6 Sol $4/1M in · index 60.9GPT-4.1 nano $0.1/1M in · index 9.6o3-mini $1.1/1M in · index 19.2GPT-5.5 $5/1M in · index 56.3GPT-5 $1.25/1M in · index 35.3GPT-5.4 $2.5/1M in · index 53.1GPT-5.2 Codex $1.75/1M in · index 41.2GPT-5.4 nano $0.2/1M in · index 39.7GPT-5.4 mini $0.75/1M in · index 40.9GPT-5.6 Luna $0.2/1M in · index 52.3GPT-5.2 $1.75/1M in · index 43.3GPT-5 Mini $0.25/1M in · index 25.8GPT-5.1 $1.25/1M in · index 37.5GPT-5.1 Codex $1.25/1M in · index 35.6o3 $2/1M in · index 31.1GPT-5.6 Terra $2/1M in · index 56.6GPT-4.1 $2/1M in · index 19.6Gemini 3.5 Flash $1.5/1M in · index 52.0Gemini 2.5 Flash $0.3/1M in · index 20.3Gemini 3.5 Flash Lite $0.3/1M in · index 37.4Gemini 3.1 Flash Lite Preview $0.25/1M in · index 25.6Gemma 4 26B A4B IT $0.07/1M in · index 26.1Gemini 3.6 Flash $0.75/1M in · index 51.6Gemma 4 31B IT $0.09/1M in · index 29.7Gemini 3.1 Pro Preview $2/1M in · index 47.7Gemini 2.5 Pro $1.25/1M in · index 25.9Gemini 2.5 Flash-Lite $0.1/1M in · index 6.7Grok 4.3 $1.25/1M in · index 37.9Grok 4.5 $2/1M in · index 55.8Grok Build 0.1 $1/1M in · index 40.7Muse Spark 1.1 $1.25/1M in · index 53.2Devstral 2 $0.4/1M in · index 19.2Magistral Small $0.5/1M in · index 10.6Qwen3-Coder 480B-A35B Instruct $1.5/1M in · index 18.2Qwen3.7 Plus $0.5/1M in · index 39.4Qwen3 32B $0.7/1M in · index 11.4Qwen3.6 35B-A3B $0.248/1M in · index 32.1Qwen3.5 27B $0.3/1M in · index 34.6Qwen3.7 Max $2.5/1M in · index 46.7Qwen3-Next 80B-A3B (Thinking) $0.5/1M in · index 16.9Qwen3.6 27B $0.6/1M in · index 37.7Qwen3.5 35B-A3B $0.25/1M in · index 29.9Qwen3 14B $0.35/1M in · index 10.4Qwen2.5 32B Instruct $0.7/1M in · index 7.2Qwen3 235B-A22B $0.7/1M in · index 19.9Qwen3 8B $0.18/1M in · index 8.3Qwen Turbo $0.05/1M in · index 6.0Qwen3.5 397B-A17B $0.6/1M in · index 34.3Qwen3.6 Plus $0.5/1M in · index 40.5Qwen3.5 122B-A10B $0.4/1M in · index 32.8Qwen3-Next 80B-A3B Instruct $0.5/1M in · index 13.8Kimi K2.5 $0.6/1M in · index 36.0Kimi K2.6 $0.95/1M in · index 45.1Kimi K2.7 Code $0.95/1M in · index 43.0Kimi K2 Thinking $0.6/1M in · index 17.2Kimi K3 $3/1M in · index 59.7GLM-5 $1/1M in · index 40.6GLM-5.1 $1.4/1M in · index 41.0GLM-5.2 $1.4/1M in · index 52.6GLM-5V-Turbo $5/1M in · index 35.3GLM-4.7 $0.6/1M in · index 34.5GLM-4.5V $0.6/1M in · index 9.0GLM-4.5 $0.6/1M in · index 19.7GLM-4.6 $0.6/1M in · index 29.3GLM-4.6V $0.3/1M in · index 10.9MiniMax-M2.7 $0.3/1M in · index 38.9MiniMax-M2.1 $0.3/1M in · index 32.1MiniMax-M2.5 $0.3/1M in · index 34.5MiniMax-M3 $0.3/1M in · index 45.4Nemotron 3 Super $0.2/1M in · index 25.7Nemotron 3 Ultra 550B A55B $0.5/1M in · index 38.3Sonar $1/1M in · index 11.6Sonar Pro $3/1M in · index 9.1Sonar Reasoning Pro $2/1M in · index 18.0Mercury 2 $0.25/1M in · index 21.9Step 3.7 Flash $0.185/1M in · index 30.9Claude Opus 5 $5/1M in · index 63.1Qwen3.8 Max $2/1M in · index 53.4Muse Spark 1.2 $1.25/1M in · index 56.8Grok 4.6 $2/1M in · index 60.9Gemini 3.7 Flash $0.75/1M in · index 56.0GLM-5.3 $1.4/1M in · index 59.5GLM-5.3-Flash $0.075/1M in · index 57.5Claude Fable 5.1 $10/1M in · index 65.7Gemini 3.8 Flash $0.75/1M in · index 58.7Claude Fable 5.1Claude Opus 5Claude Fable 5GPT-5.6 SolGrok 4.6Kimi K3

New models

How to read the numbers

  • independent — measured by an independent evaluator
  • crowd — human preference votes (Elo)
  • mirror — mirrored via an aggregator API
  • vendor — self-reported by the model's vendor

Every score keeps a link to where it came from and when we saw it. Read the methodology.