BenchLeader
xAIReasoning modelAuto-detected

Grok 4.6

Best configuration ranks #42 of 610 on the BenchLeader Index at 63.8 ±6.5 (medium reasoning effort). Last measured 12 Aug 2026. Released 12 Aug 2026.

Blended price
$3.00/M
$2.00 in · $6.00 out
Output speed
55 tok/s
First answer
37 s
first token 1.31 s
Context
500k
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Overall index64
  2. Reasoning62
  3. Coding64
  4. Knowledge77
  5. Long context67
  6. Composite83

Reasoning-effort configurations

The same model behaves differently depending on how much it is allowed to think. Each row is one setting, scored only on the benchmarks that were run at that setting. “Default” means the publisher did not say which setting was used.

EffortIndexRankSpeedFirst answerChat reply costCategories
low59.6#9957 tok/s7.71 s$0.0026Coding 55 · Composite 74 · Knowledge 76 · Long context 67 · Reasoning 50
mediumbest63.8#4255 tok/s37 s$0.0026Coding 64 · Composite 83 · Knowledge 77 · Long context 67 · Reasoning 62
high58.7#11855 tok/s41 s$0.0029Agents & tools 45 · Coding 66 · Knowledge 60 · Maths 67 · Reasoning 65
xhigh62.3#5655 tok/s48 s$0.0026Agents & tools 55 · Coding 60 · Composite 85 · Knowledge 68 · Long context 67 · Maths 61 · Reasoning 61
default63.6#4455 tok/s41 s$0.0026Agents & tools 69 · Coding 65 · Composite 73 · Knowledge 79 · Long context 67 · Maths 59 · Reasoning 71

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

BenchmarklowmediumhighxhighdefaultSourceTrend
GPQA Diamond94.0%#893.2%#16Epoch AI Benchmarking Hub
SimpleBench75.9%#8SimpleBench
LiveBench Reasoningnot in index90.5%#7LiveBench
GPQA Diamond (AA)not in index87.9%#9293.5%#1493.5%#1495.0%#4Artificial Analysis
Humanity's Last Exam (AA)not in index27.6%#12642.1%#4444.1%#3042.9%#34Artificial Analysis
GPQA Diamond (Vals)not in index94.7%#3Vals AI
ARC-AGI-174.8%#9387.5%#5987.0%#6387.0%#63ARC Prize
ARC-AGI-227.6%#9161.3%#5465.1%#4767.1%#43ARC Prize
ARC-AGI-32.1%#17ARC Prize

Coding

BenchmarklowmediumhighxhighdefaultSourceTrend
SciCode48.4%#7154.6%#2853.6%#3751.6%#49SciCode
WeirdML67.3%#34WeirdML
FrontierCode48.0%#5Cognition
LMArena WebDev1624#11LMArena
LiveBench Codingnot in index76.8%#31LiveBench
SciCode (AA)not in index49.4%#7555.9%#2653.0%#5056.5%#20Artificial Analysis
LiveCodeBench88.2%#7Vals AI
SWE-bench (Vals)not in index95.6%#4Vals AI

Agents & tools

BenchmarklowmediumhighxhighdefaultSourceTrend
Terminal-Bench20.3%#55Terminal-Bench
APEX-Agents41.2%#8Mercor
LMArena Agent3.2#16LMArena
LiveBench Agentic Codingnot in index57.0%#18LiveBench
Terminal-Bench 2.1 (Vals)78.3%#10Vals AI

Maths

BenchmarklowmediumhighxhighdefaultSourceTrend
FrontierMath Tiers 1–366.0%#21Epoch AI Benchmarking Hub
FrontierMath Tier 431.7%#25Epoch AI Benchmarking Hub
OTIS Mock AIME97.8%#1799.2%#10Epoch AI Benchmarking Hub
ProofBench51.0%#16Vals AI
LiveBench Mathematicsnot in index92.6%#16LiveBench

Knowledge

BenchmarklowmediumhighxhighdefaultSourceTrend
SimpleQA Verified49.3%#2348.9%#24Epoch AI Benchmarking Hub
LiveBench Data Analysisnot in index73.9%#30LiveBench
AA-Omniscience25.9#2928#2429.3#1930.5#17Artificial Analysis
MMLU-Pro89.4%#11Vals AI
LegalBench86.3%#12Vals AI
CorpFin66.2%#20Vals AI
TaxEval71.1%#82Vals AI

Instruction following

BenchmarklowmediumhighxhighdefaultSourceTrend
LiveBench Languagenot in index83.7%#13LiveBench

Long context

BenchmarklowmediumhighxhighdefaultSourceTrend
AA-LCR80.7%#4181.0%#3781.0%#3780.3%#45Artificial Analysis

Composite

BenchmarklowmediumhighxhighdefaultSourceTrend
Epoch Capabilities Indexnot in index156.3#16Epoch AI Benchmarking Hub
LiveBench78.0%#11LiveBench
AA Intelligence Index35.4#5143.0#2544.3#2244.4#21Artificial Analysis
Vals Indexnot in index59.2#14Vals AI

Where to run it

Every provider serving this model through OpenRouter, with throughput and first-token latency measured on live traffic over the last 30 minutes and each provider’s own price. Purple marks the best in each column.

ProviderSpeedFirst tokenInput $/MOutput $/MContextQuantisation
SpaceXAI (ZDR)52 tok/s1.36 s$2.00$6.00500k
SpaceXAI52 tok/s0.96 s$2.00$6.00500k
SpaceXAI Priority (ZDR)52 tok/s1.31 s$4.00$12.00500k
SpaceXAI Priority50 tok/s1.19 s$4.00$12.00500k
Amazon Bedrock (US)17 tok/s24 s$2.20$6.60500k

What a task costs

Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache at $0.550 per 1M. Reasoning tokens are not modelled.

WorkloadTokens in / outCostWith cachingTime
Chat reply400 / 300$0.0026$0.002242.2 s
Summarise a 30-page report12,000 / 600$0.028$0.01547.6 s
Code edit6,000 / 1,500$0.021$0.0151.1 min
Agentic coding session60,000 / 4,000$0.144$0.0791.8 min
Structured extraction2,000 / 200$0.0052$0.003040.3 s

See also

Data as of 9 Sept 2026. Compare these configurations.