llm.ing

Methodology

llm.ing runs no benchmarks of its own. It aggregates results that reputable sources publish, keeps the origin of every number attached to it, and refuses to blur the line between an independent measurement and a vendor's marketing claim.

Sources

  • models.dev (MIT) — the canonical model registry: specs, context windows, modalities, and first-party pricing. Community-maintained by SST.
  • OpenRouter — live routed pricing (including cache tiers), per-provider uptime, latency and throughput, plus mirrored Artificial Analysis indices and Design Arena ratings from its public models API. We use only documented API endpoints.
  • LMArena — crowd-voted Elo across text, web-dev, vision and agent categories, from the officialleaderboard dataset(CC BY 4.0).
  • Artificial Analysis — independently-run intelligence, coding and agentic indices plus speed measurements. Displayed with attribution; excluded from our public API per their terms.
  • A small curated seed list covers models that matter on leaderboards but are missing from registries (marked as curated in their descriptions).

Provenance labels

  • independent — measured by an evaluator with no stake in the result (e.g. Artificial Analysis running its own harness).
  • crowd — aggregated human preference votes (Elo with confidence intervals).
  • mirror — a number relayed through an aggregator's API rather than fetched from the original evaluator.
  • vendor — self-reported by the model's maker in a launch post or model card. Useful on day one, never treated as neutral.

When several sources report the same benchmark, we display the value with the strongest provenance and keep the rest visible on the model page.

Model identity & variants

Every source names models differently — dated snapshots, reasoning-effort suffixes, agent scaffolds. We keep a canonical registry keyed by the vendor's API id and map each source's names onto it. Reasoning effort and thinking budgets are tracked as variants of a model, not separate models, and a leaderboard cell shows the best variant with its label attached. Names we cannot confidently map are queued for review — never silently dropped, never guessed into the wrong row.

Freshness

Ingestion runs every six hours. Each source's last successful sync is shown on thestatus page, and score history is append-only: when a number changes we keep the old observation, so a model's record can be replayed over time.

Licensing of this site's data

The public JSON API exposes model specs, pricing, and scores whose licenses permit redistribution (e.g. LMArena's CC BY 4.0 data, with attribution). Artificial Analysis numbers are shown on this site with credit but are not included in API responses — get them from the source.