Every LLM benchmark, one honest table.
Scores, prices, and speed for 151 models — aggregated from public sources, with the origin and provenance of every number visible. Refreshed automatically, last 6h ago.
Frontier right now
Full leaderboard →| # | Model | Lab | Context | $/1M in · out | Intelligence | Coding | Agentic | Arena Elo |
|---|---|---|---|---|---|---|---|---|
| 1 | Claude Fable 5.1 | Anthropic | 1M | $10 · $50 | 65.7 | 81.6 | 61.3 | — |
| 2 | Claude Opus 5 | Anthropic | 1M | $5 · $25 | 63.1 | 78.0 | 59.2 | 1504 |
| 3 | Claude Fable 5 | Anthropic | 1M | $10 · $50 | 62.1 | 76.5 | 56.6 | 1494 |
| 4 | GPT-5.6 Sol | OpenAI | 1.1M | $4 · $20 | 60.9 | 77.4 | 57.8 | 1455 |
| 5 | Grok 4.6 | xAI | 500K | $2 · $6 | 60.9 | 76.8 | 58.7 | 1444 |
| 6 | Kimi K3open | Moonshot AI | 1.0M | $3 · $15 | 59.7 | 76.2 | 54.3 | 1482 |
| 7 | GLM-5.3open | Zhipu AI | 1M | $1.4 · $4.4 | 59.5 | 74.8 | 59.1 | — |
| 8 | Gemini 3.8 Flash | 1.0M | $0.75 · $3.75 | 58.7 | 76.3 | 50.0 | — | |
| 9 | GLM-5.3-Flash | Zhipu AI | 1M | $0.075 · $0.25 | 57.5 | 71.5 | 58.2 | 1471 |
| 10 | Claude Opus 4.8 | Anthropic | 1M | $5 · $25 | 57.3 | 74.3 | 49.4 | 1462 |
| 11 | Muse Spark 1.2 | Meta | 1.0M | $1.25 · $4.25 | 56.8 | 72.2 | 49.3 | — |
| 12 | GPT-5.6 Terra | OpenAI | 1.1M | $2 · $12 | 56.6 | 76.7 | 50.2 | 1447 |
| 13 | GPT-5.5 | OpenAI | 1.1M | $5 · $30 | 56.3 | 74.9 | 47.4 | 1472 |
| 14 | Gemini 3.7 Flash | 1.0M | $0.75 · $3.75 | 56.0 | 76.1 | 45.1 | 1491 | |
| 15 | Grok 4.5 | xAI | 500K | $2 · $6 | 55.8 | 72.4 | 48.9 | 1452 |
Intelligence, coding & agentic indices by Artificial Analysis; Arena Elo by LMArena (CC BY 4.0). Best value across model variants shown.
Intelligence vs. price
The frontier you actually pay for — up and to the left is better. Hover a point for details; click it to open the model. Free-tier models excluded.
API-onlyOpen weights
New models
- Muse Spark 1.3 ContributorMetaSep 2, 2026
- Muse Spark 1.3MetaSep 2, 2026
- Gemini 3.8 FlashGoogleSep 2, 2026
- Claude Fable 5.1AnthropicSep 1, 2026
- Qwen3.8 FlashAlibabaAug 26, 2026
- GLM-5.3-FlashZhipu AIAug 26, 2026
- DeepSeek V4 Flash Vision ExpDeepSeekAug 21, 2026
- GLM-5.3Zhipu AIAug 14, 2026
How to read the numbers
- independent — measured by an independent evaluator
- crowd — human preference votes (Elo)
- mirror — mirrored via an aggregator API
- vendor — self-reported by the model's vendor
Every score keeps a link to where it came from and when we saw it. Read the methodology.