BenchLeader

Grok 4.7 vs Qwen3.8 2.4T A95B

Verdict
  • Grok 4.7 (high) and Qwen3.8 2.4T A95B are level on quality (65.8 vs 65.1).
  • Grok 4.7 (high) is stronger in composite, knowledge.
  • Qwen3.8 2.4T A95B is stronger in long context, reasoning, agents & tools.
  • They cost about the same ($3.00 per 1M blended).
  • Grok 4.7 (high) streams 1.4× faster (55 vs 40 tokens per second).
MetricGrok 4.7 (high)Qwen3.8 2.4T A95B
BenchLeader Index65.865.1
Composite score87.579.4
Knowledge score72.666.1
Long context score64.866.6
Reasoning score76.480.2
Agents & tools score78.3
Blended price $/M$3.00$3.00
Output speed55 tok/s40 tok/s
Time to first answer1.4 s53.1 s
Context window500k984k
AA Intelligence Index46.339.9
AA-LCR77.0%80.3%
AA-Omniscience30.94.3
GPQA Diamond (AA)93.5%
Humanity's Last Exam (AA)42.3%42.5%
SciCode (AA)57.8%54.0%
LegalBench85.1%
CritPt18.0%20.0%
GDPval (AA)59.7%54.9%
τ²-Bench Banking (AA)49.1%
Code Migration34.7%
Excel Modeling Benchmark54.9%
Harvey's Legal Agent Benchmark19.6%
Legal Research Bench39.9%
MedCode48.7%
MedScribe87.2%
ProgramBench0.0%
Public Benefits Bench68.5%
SAGE31.0%

Data as of 2026-09-21. Best configuration of each model; every score links to its source on the model pages.

Grok 4.7 vs Qwen3.8 2.4T A95B: questions

Is Grok 4.7 better than Qwen3.8 2.4T A95B?
Grok 4.7 (high) and Qwen3.8 2.4T A95B are level on quality (65.8 vs 65.1). The BenchLeader Index combines every independent quality benchmark; Grok 4.7 (high) is ahead overall as of 2026-09-21, but check the category scores for your use.
Which is cheaper, Grok 4.7 or Qwen3.8 2.4T A95B?
Grok 4.7 is cheaper: $3.00 against $3.00 per million tokens, blended at three input tokens per output token.
Which is faster, Grok 4.7 or Qwen3.8 2.4T A95B?
Grok 4.7 streams faster: 55 against 40 output tokens per second.
Which has the larger context window?
Qwen3.8 2.4T A95B accepts more context: 984k against 500k tokens.