BenchLeader
xAIReasoning model

Grok 4 Fast

Best configuration ranks #234 of 610 on the BenchLeader Index at 53.0 ±4.2 (thinking reasoning effort). Last measured 8 Sept 2026. Released 19 Sept 2025.

Blended price
$1.56/M
$1.25 in · $2.50 out
Output speed
First answer
Context
2M
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Overall index53
  2. Reasoning59
  3. Coding42
  4. Agents & tools50
  5. Maths58
  6. Knowledge54
  7. Instruction following52
  8. Human preference59
  9. Multimodal46
  10. Long context63
  11. Composite52

Reasoning-effort configurations

The same model behaves differently depending on how much it is allowed to think. Each row is one setting, scored only on the benchmarks that were run at that setting. “Default” means the publisher did not say which setting was used.

EffortIndexRankSpeedFirst answerChat reply costCategories
no reasoning41.9#473$0.0013Coding 32 · Knowledge 43 · Maths 42
thinkingbest53.0#234$0.0013Agents & tools 50 · Coding 42 · Composite 52 · Human preference 59 · Instruction following 52 · Knowledge 54 · Long context 63 · Maths 58 · Multimodal 46 · Reasoning 59
default46.5#368$0.0002Agents & tools 47 · Coding 52 · Composite 43 · Human preference 61 · Instruction following 40 · Knowledge 38 · Long context 50 · Multimodal 32 · Reasoning 47

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

Benchmarkno reasoningthinkingdefaultSourceTrend
LMArena Hard Prompts1414#1421431#119LMArena
GPQA Diamond (AA)not in index84.8%#13260.6%#352Artificial Analysis
Humanity's Last Exam (AA)not in index19.1%#1674.5%#414Artificial Analysis
GPQA Diamond (Vals)not in index62.1%#10585.3%#47Vals AI
Kagi LLM Benchmark66.1%#3835.4%#116Kagi LLM Benchmark

Coding

Benchmarkno reasoningthinkingdefaultSourceTrend
WeirdML42.9%#89WeirdML
LMArena Coding1438#1411458#117LMArena
LMArena WebDev1162#117LMArena
LiveCodeBench46.1%#11579.0%#73Vals AI
IOI3.8%#4911.5%#32Vals AI
SWE-bench (Vals)not in index45.4%#80Vals AI

Agents & tools

Benchmarkno reasoningthinkingdefaultSourceTrend
Cybench30.0%#9Cybench
Terminal-Bench Hard18.9%#16112.1%#202Artificial Analysis
τ²-Bench Telecom (AA)not in index65.8%#16163.7%#167Artificial Analysis

Maths

Benchmarkno reasoningthinkingdefaultSourceTrend
AIME (Vals)33.3%#7091.3%#25Vals AI
MGSM88.0%#5290.9%#32Vals AI

Knowledge

Benchmarkno reasoningthinkingdefaultSourceTrend
AA-Omniscience-29.9#216-54.4#354Artificial Analysis
MMLU-Pro70.3%#11379.7%#90Vals AI
LegalBench78.4%#9380.6%#77Vals AI
CorpFin58.4%#8166.9%#15Vals AI
TaxEval71.6%#7575.7%#13Vals AI
MedQA75.4%#7692.1%#35Vals AI

Instruction following

Benchmarkno reasoningthinkingdefaultSourceTrend
IFBench50.5%#16937.7%#276Artificial Analysis

Human preference

Benchmarkno reasoningthinkingdefaultSourceTrend
LMArena Text1405#1261418#107LMArena

Multimodal

Benchmarkno reasoningthinkingdefaultSourceTrend
MMMU-Pro61.8%#17148.1%#212Artificial Analysis

Long context

Benchmarkno reasoningthinkingdefaultSourceTrend
Fiction.LiveBench 120k75.0%#7Fiction.live
AA-LCR73.7%#12524.0%#344Artificial Analysis

Composite

Benchmarkno reasoningthinkingdefaultSourceTrend
Epoch Capabilities Indexnot in index144.2#75Epoch AI Benchmarking Hub
AA Intelligence Index17.9#19411.1#296Artificial Analysis

What a task costs

Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache. Reasoning tokens are not modelled.

WorkloadTokens in / outCostWith cachingTime
Chat reply400 / 300$0.0013
Summarise a 30-page report12,000 / 600$0.017
Code edit6,000 / 1,500$0.011
Agentic coding session60,000 / 4,000$0.085
Structured extraction2,000 / 200$0.0030

See also

Data as of 9 Sept 2026. Compare these configurations.