BenchLeader

Grok 3

Best configuration ranks #295 of 610 on the BenchLeader Index at 50.0 ±3.9. Last measured 2 Sept 2026. Released 17 Feb 2025.

Blended price
$1.56/M
$1.25 in · $2.50 out
Output speed
First answer
Context
131k
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Overall index50
  2. Reasoning44
  3. Coding46
  4. Agents & tools34
  5. Maths54
  6. Knowledge52
  7. Instruction following48
  8. Human preference60
  9. Multimodal65
  10. Long context53
  11. Composite44

Reasoning-effort configurations

The same model behaves differently depending on how much it is allowed to think. Each row is one setting, scored only on the benchmarks that were run at that setting. “Default” means the publisher did not say which setting was used.

EffortIndexRankSpeedFirst answerChat reply costCategories
thinking$0.0076Composite 42
defaultbest50.0#295$0.0013Agents & tools 34 · Coding 46 · Composite 44 · Human preference 60 · Instruction following 48 · Knowledge 52 · Long context 53 · Maths 54 · Multimodal 65 · Reasoning 44

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

Coding

BenchmarkthinkingdefaultSource
WeirdML37.2%#116WeirdML
LMArena Coding1443#135LMArena
LiveCodeBench52.9%#110Vals AI
Aider Polyglot53.3%#21Aider polyglot leaderboard

Agents & tools

BenchmarkthinkingdefaultSource
APEX-Agents2.1%#60Mercor
Terminal-Bench Hard11.4%#207Artificial Analysis
τ²-Bench Telecom (AA)not in index48.8%#194Artificial Analysis

Maths

BenchmarkthinkingdefaultSource
OTIS Mock AIME55.6%#157Epoch AI Benchmarking Hub
MATH Level 588.8%#22Epoch AI Benchmarking Hub
AIME (Vals)58.8%#57Vals AI
MGSM91.3%#27Vals AI

Knowledge

BenchmarkthinkingdefaultSource
AA-Omniscience-33.2#226Artificial Analysis
MMLU-Pro80.0%#88Vals AI
LegalBench82.6%#60Vals AI
CorpFin59.7%#72Vals AI
TaxEval75.9%#10Vals AI
MedQA83.8%#61Vals AI

Instruction following

BenchmarkthinkingdefaultSource
IFBench46.9%#189Artificial Analysis

Human preference

BenchmarkthinkingdefaultSource
LMArena Text1412#119LMArena

Multimodal

BenchmarkthinkingdefaultSource
MMMU (validation)78.0%#10MMMU

Long context

BenchmarkthinkingdefaultSource
Fiction.LiveBench 120k58.3%#19Fiction.live
AA-LCR58.0%#226Artificial Analysis

Composite

BenchmarkthinkingdefaultSource
Epoch Capabilities Indexnot in index138.3#105Epoch AI Benchmarking Hub
AA Intelligence Index10.4#30512.1#270Artificial Analysis

What a task costs

Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache. Reasoning tokens are not modelled.

WorkloadTokens in / outCostWith cachingTime
Chat reply400 / 300$0.0013
Summarise a 30-page report12,000 / 600$0.017
Code edit6,000 / 1,500$0.011
Agentic coding session60,000 / 4,000$0.085
Structured extraction2,000 / 200$0.0030

See also

Data as of 9 Sept 2026. Compare these configurations.