Introducing BenchLeader: every AI benchmark, one leaderboard
Why we built an aggregator instead of another benchmark, how the BenchLeader Index works, and what you can do with the data.
Every week brings a new model and, with it, a chart in a launch post showing it beating the field. Every week the independent boards tell a subtler story: strong here, average there, expensive, slow, or simply not evaluated yet. Reading that story used to mean twelve browser tabs.
BenchLeader is the single page we wanted. It fetches the published results of the independent evaluators and arenas we trust, reconciles the dozen different names each of them uses for the same model, and lays everything out side by side: quality, speed, latency, price, and what a real task would actually cost.
What BenchLeader is not
It is not a benchmark. We run no models and grade no answers. That is deliberate. Running evaluations well is a full-time job that a handful of organisations already do far better than a side project could, and the field does not need a thirteenth set of GPQA numbers. It needs the twelve existing sets in one place, with the names lined up and the links intact.
It is also not a place for self-reported scores. A result appears here only when an independent evaluator or a crowdsourced arena published it.
The name
Every benchmark, one leaderboard. LMArena tells you what people prefer in conversation; Epoch AI's FrontierMath tells you whether a model can do research-level maths; SWE-bench tells you whether it can fix a real bug; OpenRouter tells you how fast it streams and what it costs to run. None of them alone tells you where a model stands. Read together, on one board, they do.
The BenchLeader Index
Benchmarks report on incompatible scales, so we standardise each one across the models it evaluated: the average model scores 50, and one standard deviation is 15 points. Standardised scores are averaged within categories (reasoning, coding, maths, agents and tools, knowledge, human preference, multimodal, long context) and the index is the plain average of those category scores. A model needs at least five results across two categories to be ranked, and the coverage column is always beside the index so you can see how much evidence sits behind a number.
The methodology page spells out every weight and every limitation. The short version: the index is for orientation; the benchmark tables are the evidence.
Per-task cost and time
A price per million tokens is an abstraction. "Summarise a 30-page report" is a budget line. We take each model's list price, output speed and time to first token and price five representative workloads with them, from a short chat reply to a long agentic coding session. It changes which models look cheap.
Updated by itself
The whole pipeline runs every morning: fetch, reconcile, aggregate, diff, publish. New models appear the day a leaderboard adds them, tagged as auto-detected until we have checked their metadata. Each day's snapshot is versioned, which is what draws the trend lines and writes the daily digest. If a source is down, its last good values are kept and marked stale on the sources page rather than vanishing.