BenchLeader

GPT-5 Pro vs Qwen3 6

Verdict
  • Qwen3 6 (max) leads on quality: 63.3 vs 60.8.
  • GPT-5 Pro is stronger in knowledge, multimodal.
  • Qwen3 6 (max) is stronger in reasoning, agents & tools, coding, composite, human preference, instruction following, long context, maths.
  • Qwen3 6 (max) is 14× cheaper ($2.96 vs $41.25 per 1M blended).
MetricGPT-5 ProQwen3 6 (max)
BenchLeader Index60.863.3
Knowledge score68.064.0
Multimodal score67.4
Reasoning score57.663.2
Agents & tools score72.0
Coding score58.2
Composite score64.9
Human preference score65.9
Instruction following score74.2
Long context score66.9
Maths score64.2
Blended price $/M$41.25$2.96
Output speed59 tok/s
Time to first answer38.5 s
Context window400k246k
GPQA Diamond87.4%
OTIS Mock AIME91.1%
SWE-bench Verified (Epoch)76.7%
SimpleQA Verified52.0%
Humanity's Last Exam31.6%
SimpleBench61.6%63.0%
Epoch Capabilities Index150.3149.3
LMArena Text1460
LMArena Hard Prompts1482
LMArena Coding1509
LMArena WebDev1479
AA Intelligence Index28.4
IFBench76.6%
AA-LCR80.7%
AA-Omniscience9.2
Terminal-Bench Hard43.9%
GPQA Diamond (AA)88.8%
Humanity's Last Exam (AA)30.8%
τ²-Bench Telecom (AA)95.9%
CorpFin66.5%
SWE-bench (Vals)72.8%
PRBench Finance51.1%
PRBench Legal49.9%
VISTA52.4%
MultiNRC65.2%
Kagi LLM Benchmark76.8%
ARC-AGI-170.2%
ARC-AGI-218.3%

Data as of 2026-09-10. Best configuration of each model; every score links to its source on the model pages.

GPT-5 Pro vs Qwen3 6: questions

Is GPT-5 Pro better than Qwen3 6?
Qwen3 6 (max) leads on quality: 63.3 vs 60.8. The BenchLeader Index combines every independent quality benchmark; Qwen3 6 (max) is ahead overall as of 2026-09-10, but check the category scores for your use.
Which is cheaper, GPT-5 Pro or Qwen3 6?
Qwen3 6 is cheaper: $2.96 against $41.25 per million tokens, blended at three input tokens per output token.
Which has the larger context window?
GPT-5 Pro accepts more context: 400k against 246k tokens.