BenchLeader

Grok 4.3 vs Qwen3 8

Verdict
  • Qwen3 8 (max) leads on quality: 63.6 vs 60.7.
  • Grok 4.3 (medium) is stronger in agents & tools, instruction following, knowledge.
  • Qwen3 8 (max) is stronger in composite, long context, multimodal, coding, human preference, maths, reasoning.
  • Grok 4.3 (medium) is 1.7× cheaper ($1.56 vs $2.67 per 1M blended).
  • Grok 4.3 (medium) streams 3.0× faster (112 vs 38 tokens per second).
MetricGrok 4.3 (medium)Qwen3 8 (max)
BenchLeader Index60.763.6
Agents & tools score60.156.0
Composite score60.370.9
Instruction following score80.0
Knowledge score72.162.3
Long context score64.065.7
Multimodal score60.267.2
Coding score69.7
Human preference score68.3
Maths score62.4
Reasoning score68.1
Blended price $/M$1.56$2.67
Output speed112 tok/s38 tok/s
Time to first answer12.1 s55.3 s
Context window1M1M
SciCode52.9%
ProofBench58.0%
Epoch Capabilities Index156.6
LMArena Text1480
LMArena Hard Prompts1502
LMArena Coding1520
LMArena WebDev1670
LMArena Vision1313
LMArena Agent3.9
LiveBench78.5%
LiveBench Reasoning88.2%
LiveBench Coding72.9%
LiveBench Agentic Coding64.7%
LiveBench Mathematics91.3%
LiveBench Data Analysis78.4%
LiveBench Language79.7%
AA Intelligence Index24.840.3
IFBench83.3%
AA-LCR75.0%78.3%
MMMU-Pro75.8%82.3%
AA-Omniscience16.73.4
Terminal-Bench Hard30.3%
GPQA Diamond (AA)89.0%92.7%
Humanity's Last Exam (AA)30.0%43.0%
SciCode (AA)53.2%
τ²-Bench Telecom (AA)91.2%
LiveCodeBench87.8%
MMLU-Pro88.6%
IOI73.0%
LegalBench83.6%
CorpFin65.8%
TaxEval75.5%
Terminal-Bench 2.1 (Vals)67.4%
SWE-bench (Vals)85.6%
GPQA Diamond (Vals)93.7%
Vals Index51.8

Data as of 2026-09-10. Best configuration of each model; every score links to its source on the model pages.

Grok 4.3 vs Qwen3 8: questions

Is Grok 4.3 better than Qwen3 8?
Qwen3 8 (max) leads on quality: 63.6 vs 60.7. The BenchLeader Index combines every independent quality benchmark; Qwen3 8 (max) is ahead overall as of 2026-09-10, but check the category scores for your use.
Is Grok 4.3 better than Qwen3 8 for agentic tasks?
Grok 4.3 scores higher in agentic tasks (60 vs 56 on the category index, where 50 is average).
Which is cheaper, Grok 4.3 or Qwen3 8?
Grok 4.3 is cheaper: $1.56 against $2.67 per million tokens, blended at three input tokens per output token.
Which is faster, Grok 4.3 or Qwen3 8?
Grok 4.3 streams faster: 112 against 38 output tokens per second.
Which has the larger context window?
Both accept 1M tokens of context.