BenchLeader

Grok 4.6 vs Qwen3 7

Verdict
  • Grok 4.6 (medium) leads on quality: 63.9 vs 61.3.
  • Grok 4.6 (medium) is stronger in coding, composite, knowledge, long context.
  • Qwen3 7 (max) is stronger in reasoning, agents & tools, human preference, instruction following, maths.
  • Grok 4.6 (medium) is 1.3× cheaper ($3.00 vs $3.75 per 1M blended).
  • Qwen3 7 (max) streams 3.2× faster (169 vs 53 tokens per second).
MetricGrok 4.6 (medium)Qwen3 7 (max)
BenchLeader Index63.961.3
Coding score64.161.4
Composite score83.456.4
Knowledge score77.563.6
Long context score67.166.0
Reasoning score62.166.8
Agents & tools score56.0
Human preference score57.2
Instruction following score77.6
Maths score57.4
Blended price $/M$3.00$3.75
Output speed53 tok/s169 tok/s
Time to first answer33.4 s16.5 s
Context window500k1M
GPQA Diamond90.9%
FrontierMath Tiers 1–364.6%
FrontierMath Tier 434.1%
OTIS Mock AIME95.6%
SWE-bench Verified (Epoch)77.3%
SimpleQA Verified55.8%
SimpleBench70.4%
SciCode54.6%48.8%
ProofBench26.0%
Epoch Capabilities Index153.7
LMArena Text1474
LMArena Hard Prompts1495
LMArena Coding1525
LMArena WebDev1517
LMArena Agent-3.1
LiveBench73.1%
LiveBench Reasoning83.3%
LiveBench Coding74.2%
LiveBench Agentic Coding43.6%
LiveBench Mathematics85.3%
LiveBench Data Analysis71.8%
LiveBench Language79.7%
AA Intelligence Index43.029.9
IFBench80.5%
AA-LCR81.0%79.0%
AA-Omniscience2813.5
Terminal-Bench Hard50.8%
GPQA Diamond (AA)93.5%92.3%
Humanity's Last Exam (AA)42.1%40.5%
SciCode (AA)55.9%49.5%
τ²-Bench Telecom (AA)94.7%
LiveCodeBench87.1%
MMLU-Pro89.3%
IOI46.8%
LegalBench84.9%
CorpFin63.7%
TaxEval75.3%
Terminal-Bench 2.1 (Vals)61.0%
SWE-bench (Vals)68.8%
GPQA Diamond (Vals)90.2%
Vals Index44.8
EQ-Bench 41110
ARC-AGI-187.5%
ARC-AGI-261.3%

Data as of 2026-09-10. Best configuration of each model; every score links to its source on the model pages.

Grok 4.6 vs Qwen3 7: questions

Is Grok 4.6 better than Qwen3 7?
Grok 4.6 (medium) leads on quality: 63.9 vs 61.3. The BenchLeader Index combines every independent quality benchmark; Grok 4.6 (medium) is ahead overall as of 2026-09-10, but check the category scores for your use.
Is Grok 4.6 better than Qwen3 7 for coding?
Grok 4.6 scores higher in coding (64 vs 61 on the category index, where 50 is average).
Which is cheaper, Grok 4.6 or Qwen3 7?
Grok 4.6 is cheaper: $3.00 against $3.75 per million tokens, blended at three input tokens per output token.
Which is faster, Grok 4.6 or Qwen3 7?
Qwen3 7 streams faster: 169 against 53 output tokens per second.
Which has the larger context window?
Qwen3 7 accepts more context: 1M against 500k tokens.