BenchLeader
xAIReasoning modelAuto-detected

Grok 4.5

Best configuration ranks #47 of 610 on the BenchLeader Index at 63.4 ±3.8. Last measured 8 Sept 2026. Released 8 Jul 2026.

Blended price
$3.00/M
$2.00 in · $6.00 out
Output speed
56 tok/s
First answer
13 s
first token 1.08 s
Context
500k
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Overall index63
  2. Reasoning69
  3. Coding60
  4. Agents & tools59
  5. Knowledge76
  6. Human preference67
  7. Multimodal65
  8. Long context66
  9. Composite66

Reasoning-effort configurations

The same model behaves differently depending on how much it is allowed to think. Each row is one setting, scored only on the benchmarks that were run at that setting. “Default” means the publisher did not say which setting was used.

EffortIndexRankSpeedFirst answerChat reply costCategories
low56 tok/s13 s$0.0026Reasoning 49
medium56 tok/s13 s$0.0026Reasoning 55
high54.2#20456 tok/s13 s$0.0026Agents & tools 37 · Coding 63 · Knowledge 60 · Maths 55 · Reasoning 59
defaultbest63.4#4756 tok/s13 s$0.0026Agents & tools 59 · Coding 60 · Composite 66 · Human preference 67 · Knowledge 76 · Long context 66 · Multimodal 65 · Reasoning 69

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

BenchmarklowmediumhighdefaultSourceTrend
GPQA Diamond93.4%#13Epoch AI Benchmarking Hub
SimpleBench70.0%#14SimpleBench
LMArena Hard Prompts1495#33LMArena
LiveBench Reasoningnot in index87.2%#22LiveBench
GPQA Diamond (AA)not in index93.1%#24Artificial Analysis
Humanity's Last Exam (AA)not in index42.7%#36Artificial Analysis
GPQA Diamond (Vals)not in index92.9%#12Vals AI
Kagi LLM Benchmark83.5%#6Kagi LLM Benchmark
ARC-AGI-179.2%#8387.2%#6185.7%#73ARC Prize
ARC-AGI-233.1%#8652.6%#7152.6%#71ARC Prize
ARC-AGI-30.3%#260.3%#260.3%#26ARC Prize

Coding

BenchmarklowmediumhighdefaultSourceTrend
SciCode54.0%#34SciCode
WeirdML46.4%#76WeirdML
FrontierCode42.4%#12Cognition
LMArena Coding1523#22LMArena
LMArena WebDev1556#23LMArena
LiveBench Codingnot in index68.6%#49LiveBench
SciCode (AA)not in index55.0%#34Artificial Analysis
LiveCodeBench87.3%#14Vals AI
SWE-bench (Vals)not in index86.6%#13Vals AI

Agents & tools

BenchmarklowmediumhighdefaultSourceTrend
Terminal-Bench12.4%#61Terminal-Bench
APEX-Agents34.2%#20Mercor
LMArena Agent3.9#14LMArena
LiveBench Agentic Codingnot in index56.5%#20LiveBench
Terminal-Bench 2.1 (Vals)67.8%#25Vals AI

Maths

BenchmarklowmediumhighdefaultSourceTrend
FrontierMath Tiers 1–357.2%#31Epoch AI Benchmarking Hub
FrontierMath Tier 424.4%#36Epoch AI Benchmarking Hub
OTIS Mock AIME97.8%#17Epoch AI Benchmarking Hub
ProofBench31.0%#28Vals AI
LiveBench Mathematicsnot in index90.8%#22LiveBench

Knowledge

BenchmarklowmediumhighdefaultSourceTrend
SimpleQA Verified48.3%#26Epoch AI Benchmarking Hub
LiveBench Data Analysisnot in index73.0%#32LiveBench
AA-Omniscience25.3#30Artificial Analysis
MMLU-Pro89.2%#14Vals AI
LegalBench86.0%#16Vals AI
CorpFin67.4%#12Vals AI
TaxEval71.7%#73Vals AI

Instruction following

BenchmarklowmediumhighdefaultSourceTrend
LiveBench Languagenot in index82.8%#16LiveBench

Human preference

BenchmarklowmediumhighdefaultSourceTrend
LMArena Text1471#38LMArena

Multimodal

BenchmarklowmediumhighdefaultSourceTrend
LMArena Vision1291#22LMArena
MMMU-Pro80.4%#34Artificial Analysis

Long context

BenchmarklowmediumhighdefaultSourceTrend
AA-LCR79.3%#65Artificial Analysis

Composite

BenchmarklowmediumhighdefaultSourceTrend
Epoch Capabilities Indexnot in index153.9#30Epoch AI Benchmarking Hub
LiveBench75.8%#26LiveBench
AA Intelligence Index39.1#41Artificial Analysis
Vals Indexnot in index51.5#27Vals AI

Where to run it

Every provider serving this model through OpenRouter, with throughput and first-token latency measured on live traffic over the last 30 minutes and each provider’s own price. Purple marks the best in each column.

ProviderSpeedFirst tokenInput $/MOutput $/MContextQuantisation
SpaceXAI (ZDR)54 tok/s0.96 s$2.00$6.00500k
SpaceXAI Priority (ZDR)50 tok/s1.44 s$4.00$12.00500k
SpaceXAI48 tok/s1.08 s$2.00$6.00500k

What a task costs

Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache at $0.300 per 1M. Reasoning tokens are not modelled.

WorkloadTokens in / outCostWith cachingTime
Chat reply400 / 300$0.0026$0.002118.1 s
Summarise a 30-page report12,000 / 600$0.028$0.01223.5 s
Code edit6,000 / 1,500$0.021$0.01339.7 s
Agentic coding session60,000 / 4,000$0.144$0.0681.4 min
Structured extraction2,000 / 200$0.0052$0.002716.3 s

See also

Data as of 9 Sept 2026. Compare these configurations.