Best configuration ranks #273 of 610 on the BenchLeader Index at 51.0 ±5.1 (thinking reasoning effort). Last measured 2 Sept 2026. Released 1 Sept 2025.
Blended price
$0.466/M
$0.144 in · $1.43 out
Output speed
190 tok/s
First answer
13 s
first token 2.13 s
Context
131k
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
Overall index51
Reasoning58
Coding56
Agents & tools42
Knowledge40
Instruction following60
Human preference55
Long context58
Composite43
Reasoning-effort configurations
The same model behaves differently depending on how much it is allowed to think. Each row is one setting, scored only on the benchmarks that were run at that setting. “Default” means the publisher did not say which setting was used.
Effort
Index
Rank
Speed
First answer
Chat reply cost
Categories
thinkingbest
51.0
#273
190 tok/s
13 s
$0.0005
Agents & tools 42 · Coding 56 · Composite 43 · Human preference 55 · Instruction following 60 · Knowledge 40 · Long context 58 · Reasoning 58
default
48.4
#324
176 tok/s
2.13 s
$0.0002
Agents & tools 40 · Coding 58 · Composite 41 · Human preference 58 · Instruction following 42 · Knowledge 36 · Long context 53 · Reasoning 55
Benchmark results
One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.
Every provider serving this model through OpenRouter, with throughput and first-token latency measured on live traffic over the last 30 minutes and each provider’s own price. Purple marks the best in each column.
Provider
Speed
First token
Input $/M
Output $/M
Context
Quantisation
Google Vertex
91 tok/s
0.45 s
$0.150
$1.20
262k
–
DeepInfra
79 tok/s
0.50 s
$0.090
$1.10
262k
fp8
Alibaba Cloud Int.
68 tok/s
0.52 s
$0.098
$0.780
131k
–
NovitaAI
68 tok/s
0.81 s
$0.150
$1.50
131k
bf16
Parasail
34 tok/s
1.00 s
$0.100
$1.10
262k
fp8
Price history
Listed price per 1M tokens over time, as recorded by OpenRouter for the provider with the longest history.
input outputnow $0.098 in · $0.780 out
What a task costs
Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache. Reasoning tokens are not modelled.