BenchLeader

Grok 4.6 vs Qwen3.8 2.4T A95B

Verdict
  • Grok 4.6 (medium) leads on quality: 66.2 vs 65.2.
  • Grok 4.6 (medium) is stronger in coding, composite, knowledge, long context.
  • Qwen3.8 2.4T A95B is stronger in agents & tools, reasoning.
  • They cost about the same ($3.00 per 1M blended).
  • Grok 4.6 (medium) streams 1.5× faster (57 vs 38 tokens per second).
MetricGrok 4.6 (medium)Qwen3.8 2.4T A95B
BenchLeader Index66.265.2
Agents & tools score73.378.3
Coding score64.1
Composite score83.679.8
Knowledge score77.666.3
Long context score67.066.6
Reasoning score67.780.5
Blended price $/M$3.00$3.00
Output speed57 tok/s38 tok/s
Time to first answer34.1 s55.2 s
Context window500k984k
SciCode54.6%
AA Intelligence Index43.040.0
AA-LCR81.0%80.3%
AA-Omniscience284.3
GPQA Diamond (AA)93.5%93.5%
Humanity's Last Exam (AA)42.1%42.5%
SciCode (AA)55.9%54.0%
ARC-AGI-187.5%
ARC-AGI-261.3%
CritPt17.7%20.0%
GDPval (AA)57.1%56.4%
τ²-Bench Banking (AA)44.3%49.1%
DeepSWE67.5%
CursorBench67.1%

Data as of 2026-09-19. Best configuration of each model; every score links to its source on the model pages.

Grok 4.6 vs Qwen3.8 2.4T A95B: questions

Is Grok 4.6 better than Qwen3.8 2.4T A95B?
Grok 4.6 (medium) leads on quality: 66.2 vs 65.2. The BenchLeader Index combines every independent quality benchmark; Grok 4.6 (medium) is ahead overall as of 2026-09-19, but check the category scores for your use.
Is Grok 4.6 better than Qwen3.8 2.4T A95B for agentic tasks?
Qwen3.8 2.4T A95B scores higher in agentic tasks (78 vs 73 on the category index, where 50 is average).
Which is cheaper, Grok 4.6 or Qwen3.8 2.4T A95B?
Grok 4.6 is cheaper: $3.00 against $3.00 per million tokens, blended at three input tokens per output token.
Which is faster, Grok 4.6 or Qwen3.8 2.4T A95B?
Grok 4.6 streams faster: 57 against 38 output tokens per second.
Which has the larger context window?
Qwen3.8 2.4T A95B accepts more context: 984k against 500k tokens.