BenchLeader
xAIReasoning modelAuto-detected

Grok 4.3

Best configuration ranks #81 of 610 on the BenchLeader Index at 60.7 ±6.0 (medium reasoning effort). Released 17 Apr 2026.

Blended price
$1.56/M
$1.25 in · $2.50 out
Output speed
113 tok/s
First answer
12 s
first token 0.51 s
Context
1M
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Overall index61
  2. Agents & tools60
  3. Knowledge72
  4. Instruction following80
  5. Multimodal60
  6. Long context64
  7. Composite60

Reasoning-effort configurations

The same model behaves differently depending on how much it is allowed to think. Each row is one setting, scored only on the benchmarks that were run at that setting. “Default” means the publisher did not say which setting was used.

EffortIndexRankSpeedFirst answerChat reply costCategories
no reasoning48.4#327113 tok/s0.69 s$0.0013Agents & tools 50 · Composite 47 · Instruction following 49 · Knowledge 48 · Long context 42 · Multimodal 49
low59.5#102114 tok/s5.06 s$0.0013Agents & tools 57 · Composite 60 · Instruction following 78 · Knowledge 71 · Long context 63 · Multimodal 57
mediumbest60.7#81113 tok/s12 s$0.0013Agents & tools 60 · Composite 60 · Instruction following 80 · Knowledge 72 · Long context 64 · Multimodal 60
high50.1#293118 tok/s21 s$0.0013Agents & tools 32 · Coding 55 · Knowledge 51 · Maths 49 · Reasoning 64
default57.8#136118 tok/s21 s$0.0013Agents & tools 67 · Coding 49 · Composite 37 · Human preference 53 · Instruction following 78 · Knowledge 65 · Long context 63 · Multimodal 60 · Reasoning 63

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

Benchmarkno reasoninglowmediumhighdefaultSourceTrend
GPQA Diamond88.8%#47Epoch AI Benchmarking Hub
LMArena Hard Prompts1457#85LMArena
LiveBench Reasoningnot in index70.8%#50LiveBench
GPQA Diamond (AA)not in index65.8%#32384.3%#14289.0%#8390.1%#64Artificial Analysis
Humanity's Last Exam (AA)not in index6.8%#31118.4%#17430.0%#10437.2%#70Artificial Analysis
GPQA Diamond (Vals)not in index91.4%#22Vals AI

Coding

Benchmarkno reasoninglowmediumhighdefaultSourceTrend
SciCode47.3%#74SciCode
WeirdML49.9%#66WeirdML
LMArena Coding1488#81LMArena
LMArena WebDev1357#89LMArena
LiveBench Codingnot in index69.9%#48LiveBench
SciCode (AA)not in index39.4%#11848.3%#78Artificial Analysis
LiveCodeBench84.5%#37Vals AI
IOI15.3%#29Vals AI
SWE-bench (Vals)not in index71.4%#55Vals AI

Agents & tools

Benchmarkno reasoninglowmediumhighdefaultSourceTrend
LiveBench Agentic Codingnot in index18.5%#52LiveBench
Terminal-Bench Hard18.9%#16126.5%#12630.3%#11037.9%#61Artificial Analysis
τ²-Bench Telecom (AA)not in index65.8%#16188.9%#7191.2%#5997.7%#9Artificial Analysis
Terminal-Bench 2.1 (Vals)42.0%#54Vals AI

Maths

Benchmarkno reasoninglowmediumhighdefaultSourceTrend
FrontierMath Tiers 1–342.8%#44Epoch AI Benchmarking Hub
FrontierMath Tier 414.6%#46Epoch AI Benchmarking Hub
OTIS Mock AIME93.3%#44Epoch AI Benchmarking Hub
ProofBench11.0%#49Vals AI
LiveBench Mathematicsnot in index84.3%#42LiveBench

Knowledge

Benchmarkno reasoninglowmediumhighdefaultSourceTrend
SimpleQA Verified33.2%#49Epoch AI Benchmarking Hub
LiveBench Data Analysisnot in index55.8%#49LiveBench
AA-Omniscience-33.1#22513.9#5716.7#4918.0#48Artificial Analysis
MMLU-Pro85.8%#51Vals AI
LegalBench84.5%#29Vals AI
CorpFin68.5%#7Vals AI
TaxEval70.8%#84Vals AI

Instruction following

Benchmarkno reasoninglowmediumhighdefaultSourceTrend
LiveBench Languagenot in index73.6%#44LiveBench
IFBench47.6%#18481.0%#683.3%#181.3%#5Artificial Analysis

Human preference

Benchmarkno reasoninglowmediumhighdefaultSourceTrend
LMArena Text1443#76LMArena
EQ-Bench 41075#24EQ-Bench

Multimodal

Benchmarkno reasoninglowmediumhighdefaultSourceTrend
LMArena Vision1228#63LMArena
MMMU-Pro64.8%#14872.8%#10075.8%#6778.1%#55Artificial Analysis

Long context

Benchmarkno reasoninglowmediumhighdefaultSourceTrend
AA-LCR32.3%#31274.0%#12075.0%#11173.0%#132Artificial Analysis

Composite

Benchmarkno reasoninglowmediumhighdefaultSourceTrend
Epoch Capabilities Indexnot in index149.2#49Epoch AI Benchmarking Hub
LiveBench62.3%#52LiveBench
AA Intelligence Index14.5#22924.3#12424.8#11925.4#114Artificial Analysis
Vals Indexnot in index24.3#46Vals AI

Where it wins

Benchmarks where this configuration ranks in the top five of every configuration measured.

Where to run it

Every provider serving this model through OpenRouter, with throughput and first-token latency measured on live traffic over the last 30 minutes and each provider’s own price. Purple marks the best in each column.

ProviderSpeedFirst tokenInput $/MOutput $/MContextQuantisation
SpaceXAI (ZDR)104 tok/s0.51 s$1.25$2.501M
SpaceXAI90 tok/s0.57 s$1.25$2.501M
SpaceXAI Priority65 tok/s0.51 s$2.50$5.001M
SpaceXAI Priority (ZDR)21 tok/s0.48 s$2.50$5.001M

What a task costs

Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache at $0.200 per 1M. Reasoning tokens are not modelled.

WorkloadTokens in / outCostWith cachingTime
Chat reply400 / 300$0.0013$0.000914.8 s
Summarise a 30-page report12,000 / 600$0.017$0.007117.5 s
Code edit6,000 / 1,500$0.011$0.006525.5 s
Agentic coding session60,000 / 4,000$0.085$0.03847.7 s
Structured extraction2,000 / 200$0.0030$0.001413.9 s

See also

Data as of 9 Sept 2026. Compare these configurations.