BenchLeader

Grok 4.3 vs Qwen3 7

Verdict
  • Grok 4.3 (medium) and Qwen3 7 (max) are level on quality (60.7 vs 61.3).
  • Grok 4.3 (medium) is stronger in agents & tools, composite, instruction following, knowledge, multimodal.
  • Qwen3 7 (max) is stronger in long context, coding, human preference, maths, reasoning.
  • Grok 4.3 (medium) is 2.4× cheaper ($1.56 vs $3.75 per 1M blended).
  • Qwen3 7 (max) streams 1.5× faster (169 vs 112 tokens per second).
MetricGrok 4.3 (medium)Qwen3 7 (max)
BenchLeader Index60.761.3
Agents & tools score60.156.0
Composite score60.356.4
Instruction following score80.077.6
Knowledge score72.163.6
Long context score64.066.0
Multimodal score60.2
Coding score61.4
Human preference score57.2
Maths score57.4
Reasoning score66.8
Blended price $/M$1.56$3.75
Output speed112 tok/s169 tok/s
Time to first answer12.1 s16.5 s
Context window1M1M
GPQA Diamond90.9%
FrontierMath Tiers 1–364.6%
FrontierMath Tier 434.1%
OTIS Mock AIME95.6%
SWE-bench Verified (Epoch)77.3%
SimpleQA Verified55.8%
SimpleBench70.4%
SciCode48.8%
ProofBench26.0%
Epoch Capabilities Index153.7
LMArena Text1474
LMArena Hard Prompts1495
LMArena Coding1525
LMArena WebDev1517
LMArena Agent-3.1
LiveBench73.1%
LiveBench Reasoning83.3%
LiveBench Coding74.2%
LiveBench Agentic Coding43.6%
LiveBench Mathematics85.3%
LiveBench Data Analysis71.8%
LiveBench Language79.7%
AA Intelligence Index24.829.9
IFBench83.3%80.5%
AA-LCR75.0%79.0%
MMMU-Pro75.8%
AA-Omniscience16.713.5
Terminal-Bench Hard30.3%50.8%
GPQA Diamond (AA)89.0%92.3%
Humanity's Last Exam (AA)30.0%40.5%
SciCode (AA)49.5%
τ²-Bench Telecom (AA)91.2%94.7%
LiveCodeBench87.1%
MMLU-Pro89.3%
IOI46.8%
LegalBench84.9%
CorpFin63.7%
TaxEval75.3%
Terminal-Bench 2.1 (Vals)61.0%
SWE-bench (Vals)68.8%
GPQA Diamond (Vals)90.2%
Vals Index44.8
EQ-Bench 41110

Data as of 2026-09-10. Best configuration of each model; every score links to its source on the model pages.

Grok 4.3 vs Qwen3 7: questions

Is Grok 4.3 better than Qwen3 7?
Grok 4.3 (medium) and Qwen3 7 (max) are level on quality (60.7 vs 61.3). The BenchLeader Index combines every independent quality benchmark; Qwen3 7 (max) is ahead overall as of 2026-09-10, but check the category scores for your use.
Is Grok 4.3 better than Qwen3 7 for agentic tasks?
Grok 4.3 scores higher in agentic tasks (60 vs 56 on the category index, where 50 is average).
Which is cheaper, Grok 4.3 or Qwen3 7?
Grok 4.3 is cheaper: $1.56 against $3.75 per million tokens, blended at three input tokens per output token.
Which is faster, Grok 4.3 or Qwen3 7?
Qwen3 7 streams faster: 169 against 112 output tokens per second.
Which has the larger context window?
Both accept 1M tokens of context.