BenchLeader
SpaceXAIReasoning modelNewReleased 21 Sept 2026

Grok 4.7

Reasoning effort

Grok 4.7 is a SpaceXAI proprietary reasoning model, released 21 Sept 2026. Its best configuration (high reasoning effort) ranks #33 of 431 on the BenchLeader Index at 65.8 ±8.6, in the upper half. It scores highest in composite (88) and lowest in long context (65). At $3.00 per million tokens blended it is among the most expensive ranked models. Output speed of 55 tokens per second puts it slower than most, with a first token in 1.4 s. It has been measured at 2 reasoning-effort settings; this summary describes the best-scoring one, and the tabs above switch between them. Last measured 21 Sept 2026.

Blended price
$3.00/M
$2.00 in · $6.00 out
Output speed
55 tok/s
OpenRouter traffic, 7-day median; not yet measured by Artificial Analysis
First answer
1.41 s
Context
500k
Full answer
Cost per run
Released
21 Sept 2026
How it scores by categoryDashed line = average model (50). One step of 15 = one standard deviation.
  1. Overall index66
  2. Reasoning76
  3. Knowledge73
  4. Long context65
  5. Composite88

Versions

SpaceXAI has shipped 9 models under this name. Each is ranked on its own results; a newer version often has fewer results so far, which holds its index nearer the average until more arrive.

ModelReleasedIndexRank
NewGrok 4.7highthis page21 Sept 202665.8#33
Grok 4.6medium11 Aug 202666.0#30
Grok 4.5high16 Jul 202660.9#90
Grok 4.3medium17 Apr 202659.8#111
Grok 4.20thinking5 Mar 202659.9#108
Grok 4.119 Nov 202557.0#166
Grok 49 Jul 202558.4#137
Grok 317 Feb 202551.0#321
Grok 212 Dec 202442.2#552

Reasoning-effort configurations

The same model behaves differently depending on how much it is allowed to think. Each row is one setting, scored only on the benchmarks that were run at that setting. “Not stated” collects results from publishers that did not say which setting they used; for a reasoning model that is usually its thinking mode, but we do not assume it. Pick a setting here or at the top of the page to see its price, speed and category scores.

EffortIndexRankSpeedFirst answerChat reply costCategories
highbest65.8#3355 tok/s1.41 s$0.0026Composite 88 · Knowledge 73 · Long context 65 · Reasoning 76
xhigh62.5#7255 tok/s1.41 s$0.0026Agents & tools 64 · Coding 41 · Composite 88 · Knowledge 79 · Long context 65 · Reasoning 76

Benchmark results

One column per reasoning effort. Rank is among every configuration of every model on that benchmark. Hover a score for the run it came from.

Reasoning

BenchmarkhighxhighSource
Humanity's Last Exam (AA)not in index42.3%#4343.1%#34Artificial Analysis
CritPt18.0%#4417.7%#46Artificial Analysis
MysteryMechanismnot in index25.7%#8Vals AI

Coding

BenchmarkhighxhighSource
SciCode (AA)not in index57.8%#1357.4%#15Artificial Analysis
IOI39.4%#21Vals AI
Code Migrationnot in index34.7%#24Vals AI
ProgramBenchnot in index0.0%#16Vals AI
Vibe Code Bench v1.1not in index75.9%#19Vals AI

Agents & tools

BenchmarkhighxhighSource
Terminal-Bench 2.1 (Vals)76.0%#14Vals AI
GDPval (AA)not in index59.7%#559.8%#4Artificial Analysis
Finance Agent v2not in index49.2%#30Vals AI
Harvey's Legal Agent Benchmarknot in index19.6%#5Vals AI
Legal Research Benchnot in index39.9%#19Vals AI
Tax Agent Benchnot in index62.2%#14Vals AI
Terminal-Bench 4.0 (Vals)not in index12.1%#16Vals AI
Terminal-Bench Sciencenot in index2.9%#17Vals AI

Knowledge

BenchmarkhighxhighSource
AA-Omniscience30.9#1832.0#15Artificial Analysis
LegalBench85.1%#24Vals AI
BioMysteryBenchnot in index68.5%#6Vals AI
Excel Modeling Benchmarknot in index54.9%#33Vals AI
MedCodenot in index48.7%#22Vals AI
MedScribenot in index87.2%#12Vals AI
Public Benefits Benchnot in index68.5%#4Vals AI

Multimodal

BenchmarkhighxhighSource
SAGEnot in index31.0%#67Vals AI

Long context

BenchmarkhighxhighSource
AA-LCR77.0%#9976.7%#101Artificial Analysis

Composite

BenchmarkhighxhighSource
AA Intelligence Index46.3#1746.5#16Artificial Analysis
Vals Indexnot in index54.1#23Vals AI

Where to run it

Every provider serving this model through OpenRouter, with throughput and first-token latency measured on live traffic over the last 30 minutes and each provider’s own price. Purple marks the best in each column.

ProviderSpeedFirst tokenInput $/MOutput $/MContextQuantisation
SpaceXAI61 tok/s1.38 s$1.60$4.80500k
SpaceXAI55 tok/s1.41 s$1.60$4.80500k
SpaceXAI50 tok/s1.30 s$3.20$9.60500k

What a task costs

Estimates from list price, output speed and time to first answer for the best configuration. “With caching” assumes three-quarters of the input is served from the prompt cache at $0.500 per 1M. Reasoning tokens are not modelled.

WorkloadTokens in / outCostWith cachingTime
Chat reply400 / 300$0.0026$0.00226.9 s
Summarise a 30-page report12,000 / 600$0.028$0.01412.3 s
Code edit6,000 / 1,500$0.021$0.01428.7 s
Agentic coding session60,000 / 4,000$0.144$0.0761.2 min
Structured extraction2,000 / 200$0.0052$0.00305.0 s

See also

Data as of 21 Sept 2026. Compare these configurations.

Cite as: BenchLeader, “Grok 4.7: benchmarks, pricing, speed and rank”, https://www.benchleader.com/models/grok-4-7, data as of 21 Sept 2026.