BenchLeader
xAIReasoning model

Grok 4

Best configuration ranks #144 of 610 on the BenchLeader Index at 57.2 ±3.0. Last measured 2 Sept 2026. Released 9 Jul 2025.

Blended price
$1.56/M
$1.25 in · $2.50 out
Output speed
First answer
Context
256k
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Overall index57
  2. Reasoning58
  3. Coding60
  4. Agents & tools55
  5. Maths53
  6. Knowledge58
  7. Instruction following54
  8. Human preference60
  9. Multimodal54
  10. Long context69
  11. Composite57

Reasoning-effort configurations

The same model behaves differently depending on how much it is allowed to think. Each row is one setting, scored only on the benchmarks that were run at that setting. “Default” means the publisher did not say which setting was used.

EffortIndexRankSpeedFirst answerChat reply costCategories
high$0.0057Agents & tools 34
thinking$0.0057Reasoning 45
defaultbest57.2#144$0.0013Agents & tools 55 · Coding 60 · Composite 57 · Human preference 60 · Instruction following 54 · Knowledge 58 · Long context 69 · Maths 53 · Multimodal 54 · Reasoning 58

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

BenchmarkhighthinkingdefaultSource
GPQA Diamond87.0%#64Epoch AI Benchmarking Hub
SimpleBench60.5%#30SimpleBench
LMArena Hard Prompts1420#128LMArena
GPQA Diamond (AA)not in index87.7%#94Artificial Analysis
Humanity's Last Exam (AA)not in index26.7%#129Artificial Analysis
GPQA Diamond (Vals)not in index88.1%#34Vals AI
Kagi LLM Benchmark73.6%#17Kagi LLM Benchmark
ARC-AGI-166.7%#10079.6%#82ARC Prize
ARC-AGI-216.0%#9829.4%#89ARC Prize

Coding

BenchmarkhighthinkingdefaultSource
WeirdML45.7%#80WeirdML
LMArena Coding1435#144LMArena
LiveCodeBench83.3%#48Vals AI
IOI26.2%#17Vals AI
SWE-bench (Vals)not in index57.8%#75Vals AI
Aider Polyglot79.6%#5Aider polyglot leaderboard

Agents & tools

BenchmarkhighthinkingdefaultSource
Terminal-Bench27.2%#47Terminal-Bench
GDPval21.1%#10OpenAI
Cybench43.0%#4Cybench
APEX-Agents15.2%#46Mercor
Terminal-Bench Hard37.9%#61Artificial Analysis
τ²-Bench Telecom (AA)not in index74.8%#134Artificial Analysis
BFCL Overall63.0%#8Berkeley Function Calling Leaderboard

Maths

BenchmarkhighthinkingdefaultSource
OTIS Mock AIME84.0%#89Epoch AI Benchmarking Hub
AIME (Vals)90.6%#27Vals AI
MGSM90.9%#31Vals AI
IMO 202521.4%#3MathArena
MathArena Apex2.1%#29MathArena

Knowledge

BenchmarkhighthinkingdefaultSource
AA-Omniscience2.1#82Artificial Analysis
MMLU-Pro85.3%#57Vals AI
LegalBench83.2%#51Vals AI
CorpFin66.0%#23Vals AI
TaxEval65.1%#112Vals AI
MedQA92.5%#31Vals AI

Instruction following

BenchmarkhighthinkingdefaultSource
IFBench53.7%#153Artificial Analysis

Human preference

BenchmarkhighthinkingdefaultSource
LMArena Text1411#121LMArena

Multimodal

BenchmarkhighthinkingdefaultSource
LMArena Vision1210#69LMArena
MMMU-Pro68.8%#128Artificial Analysis

Long context

BenchmarkhighthinkingdefaultSource
Fiction.LiveBench 120k96.9%#2Fiction.live
AA-LCR68.0%#180Artificial Analysis

Composite

BenchmarkhighthinkingdefaultSource
Epoch Capabilities Indexnot in index146.4#62Epoch AI Benchmarking Hub
AA Intelligence Index22.5#143Artificial Analysis

Where it wins

Benchmarks where this configuration ranks in the top five of every configuration measured.

What a task costs

Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache. Reasoning tokens are not modelled.

WorkloadTokens in / outCostWith cachingTime
Chat reply400 / 300$0.0013
Summarise a 30-page report12,000 / 600$0.017
Code edit6,000 / 1,500$0.011
Agentic coding session60,000 / 4,000$0.085
Structured extraction2,000 / 200$0.0030

See also

Data as of 9 Sept 2026. Compare these configurations.