LLM leaderboard 2026: who leads, and how to read the rankings
The current state of the AI model race across quality, speed and cost, with every number pulled from independent benchmarks and refreshed daily.
Every model launch comes with a chart in which the new model wins. This page is the other kind of chart: the one nobody paid for. BenchLeader lines up 610 configurations of 357 models from 104 providers, scored only by independent evaluators and public arenas, and refreshes the whole thing every morning. The numbers below are as of 9 Sept 2026.
Who leads right now
On the overall BenchLeader Index, the top of the table is GPT-6 Astra (max), followed by GPT-6 Astra (max) and GPT-6 Astra (max). The gap between first and second is usually inside the confidence band, which is the honest way to say the two are tied and the order can flip on any given morning.
Quality is one axis. Price is the one that decides what you actually deploy. Kimi K3 sits within 5 points of the leader at $6.00 per million tokens, 3.7× cheaper than GPT-6 Astra (max).
| # | Model | Blended $/M | Index | Price $/M |
|---|---|---|---|---|
| 1 | $1.50/M | 64.4 | $1.50 | |
| 2 | $1.50/M | 63.5 | $1.50 | |
| 3 | $1.13/M | 61.5 | $1.13 | |
| 4 | $0.12/M | 61.0 | $0.119 | |
| 5 | $1.50/M | 60.7 | $1.50 |
As of 9 Sept 2026. Full list.
How to read the index
The index is not a single benchmark. Each of the 60-odd benchmarks we track reports on its own scale, so every score is first standardised against the models that benchmark evaluated: the average model becomes 50, one standard deviation becomes 15. Those standardised scores are averaged inside a category (reasoning, coding, agents, maths, knowledge, and so on), and the index is the average across categories, pulled slightly toward 50 when a model has only a handful of results. The full recipe, weights included, is on the methodology page.
Two consequences are worth keeping in mind:
- A configuration is a model at one reasoning effort. GPT-6 Astra at max and GPT-6 Astra at low are different rows, because they are different products with different scores, speeds and latencies. The tables show each model's best-scoring effort by default and let you expand to all of them.
- Thin evidence is discounted. A model measured on five favourable benchmarks will not outrank one measured on twenty. The ± figure next to each index shows how much the benchmarks behind it disagree.
Quality against price
The home page plots quality against price on a log scale, and the shape of that chart has not changed in a year: a frontier of a dozen models that nobody beats on both axes, and a long tail underneath. The interesting movement is on the left of the frontier, where open-weights models keep pushing the price of "good enough" down. See best open weights for where that line sits today.
What this leaderboard does not do
BenchLeader runs no evaluations. It does not accept self-reported scores from model cards or launch posts. It does not weight a benchmark higher because a provider asked. When a source is down, yesterday's numbers stay in place and are marked stale rather than silently dropped. Every figure on every page links to the site that measured it, so you can check us.
Where to go next
- Best for coding, best for agents, fast and good
- Compare any two models, with a plain-language verdict
- Release watch: what arrived this week and who is due
- Providers: every lab, its line-up, prices and cadence
Frequently asked
How often is the leaderboard updated? Every morning. Each page shows the data date it was built from.
Why does a model appear more than once? Because it was measured at more than one reasoning effort. Each effort is its own configuration.
Why is a well-known model missing? Either no independent publisher has measured it yet, or it has fewer than five results across two categories, which is the minimum to receive an index. It still has a model page.