Leaderboards
The same models, ranked by whichever question you are actually asking. Every board is rebuilt daily from the latest data.
Straight answers
The questions people arrive with, answered from today’s data. Top five each; the board behind it is one click away.
- Best under $2/M
Highest quality with a blended price below $2 per million tokens.
- 1
Gemini 3.7 Flashmedium$1.50/M
- 2
Gemini 3.8 Flashmedium$1.50/M
- 3
Gemini 3 Flashthinking$1.13/M
- 4
GLM-5.3-Flash$0.12/M
- 5
Gemini 3.6 Flash$1.50/M
- 1
- Best under $0.50/M
Highest quality for budget workloads, below fifty cents per million tokens.
- 1
GLM-5.3-Flash$0.119/M
- 2
GPT-5.6 Lunaxhigh$0.450/M
- 3
DeepSeek V4 Flashhigh$0.168/M
- 4
Qwen3.8-Flash-Next$0.230/M
- 5
Qwen3.5 397B-A17B$0.387/M
- 1
- Fast and good
At least 100 tokens per second with an index above 55.
- 1
Gemini 3.7 Flashmedium282 tok/s
- 2
Gemini 3.8 Flashmedium273 tok/s
- 3
Muse Spark 1.2219 tok/s
- 4
Gemini 3.5 Flashmedium209 tok/s
- 5
Gemini 3 Flashthinking200 tok/s
- 1
- Quick answers
Index above 55 and the answer starts within two seconds.
- 1
Qwen3.5 Plusthinking1.1 s
- 2
Claude Sonnet 4.5high1.4 s
- 3
Kimi K2.51.6 s
- 4
DeepSeek V3.2thinking1.8 s
- 1
- Best open weights
Highest-quality models you can download and run yourself.
- 1
Kimi K366.5
- 2
GLM-5.3-Flash61.0
- 3
Kimi K2.660.6
- 4
GLM-5.160.1
- 5
DeepSeek V4 Prohigh60.0
- 1
- Cheapest 1M context
Context window of at least one million tokens, ordered by price.
- 1
nemotron-3-nano-30b-a3bthinking$0.09/M
- 2
GLM-5.3-Flash$0.12/M
- 3
Gemini 2.0 Flash-Lite$0.13/M
- 4
Qwen Plus$0.16/M
- 5
DeepSeek V4 Flashhigh$0.17/M
- 1
- Best for coding
Highest coding category score.
- 1
GPT-6 Astramax79.0
- 2
Claude Opus 574.4
- 3
GPT-5.6 Solhigh72.0
- 4
Qwen3 8max71.7
- 5
Claude Fable 570.6
- 1
- Best for agents
Highest agents-and-tools category score.
- 1
Claude Fable 574.4
- 2
GLM-4.6thinking72.3
- 3
Claude Opus 571.6
- 4
Grok 4.1thinking70.4
- 5
o370.2
- 1
Price cuts
Listed blended prices that fell by more than 5% in the last 90 days.
| Model | Before | After | Change | When |
|---|---|---|---|---|
| $1.19/M | $1.09/M | −9% | 9 Sept 2026 | |
| $0.645/M | $0.430/M | −33% | 9 Sept 2026 | |
| $1.71/M | $1.04/M | −39% | 9 Sept 2026 | |
| $2.15/M | $1.48/M | −31% | 9 Sept 2026 | |
| $1.29/M | $1.19/M | −8% | 7 Sept 2026 | |
| $0.836/M | $0.645/M | −23% | 7 Sept 2026 | |
| $1.05/M | $0.420/M | −60% | 7 Sept 2026 | |
| $0.440/M | $0.353/M | −20% | 7 Sept 2026 | |
| $0.525/M | $0.367/M | −30% | 7 Sept 2026 | |
| $1.23/M | $0.986/M | −20% | 5 Sept 2026 | |
| $0.109/M | $0.103/M | −5% | 5 Sept 2026 | |
| $0.645/M | $0.527/M | −18% | 5 Sept 2026 | |
| $1.30/M | $1.23/M | −6% | 4 Sept 2026 | |
| $1.71/M | $1.01/M | −41% | 4 Sept 2026 | |
| $2.15/M | $1.40/M | −35% | 4 Sept 2026 | |
| $2.15/M | $1.78/M | −17% | 3 Sept 2026 | |
| $0.438/M | $0.350/M | −20% | 2 Sept 2026 | |
| $0.112/M | $0.101/M | −10% | 31 Aug 2026 | |
| $1.05/M | $0.900/M | −14% | 30 Aug 2026 | |
| $1.14/M | $1.05/M | −8% | 29 Aug 2026 |
Core boards
Six ways to rank the same models. Each opens as a chart plus a searchable, sortable table.
- Quality: the BenchLeader Index
Composite of every quality benchmark, normalised per benchmark and averaged per category. Requires results on at least five benchmarks across two categories.
- Value: quality per dollar
BenchLeader Index divided by the blended price per million tokens (3 input : 1 output). Higher means more quality for the money.
- Speed: output tokens per second
Median output throughput measured across healthy providers over the last 30 minutes of live traffic (OpenRouter).
- Latency: time to first answer
Median seconds until the answer starts, including any reasoning phase (Artificial Analysis), or time to first streamed token where that is all we have.
- Cost: blended price per million tokens
First-party list price where published, blended at 3 input tokens per output token.
- Context window
Maximum context length in tokens as listed by the provider.
Per task
Prices per million tokens hide the shape of real work. These boards price and time five representative workloads from each model’s list price, output speed and time to first token.
| Workload | Input tokens | Output tokens | Boards |
|---|---|---|---|
Chat reply A short assistant turn: 400 tokens of context in, 300 tokens out. | 400 | 300 | CostTime |
Summarise a 30-page report 12,000 tokens of document in, a 600-token summary out. | 12,000 | 600 | CostTime |
Code edit 6,000 tokens of files and instructions in, a 1,500-token diff out. | 6,000 | 1,500 | CostTime |
Agentic coding session One long tool-using loop: 60,000 tokens in across turns, 4,000 tokens out. | 60,000 | 4,000 | CostTime |
Structured extraction 2,000 tokens of source text in, 200 tokens of JSON out. | 2,000 | 200 | CostTime |