BenchLeader

GPT-5.2 vs Qwen3 8

Verdict
  • Qwen3 8 (max) leads on quality: 63.6 vs 61.6.
  • GPT-5.2 is stronger in agents & tools, instruction following, long context.
  • Qwen3 8 (max) is stronger in coding, composite, human preference, knowledge, multimodal, reasoning, maths.
  • Qwen3 8 (max) is 1.8× cheaper ($2.67 vs $4.81 per 1M blended).
  • GPT-5.2 streams 1.8× faster (68 vs 38 tokens per second).
MetricGPT-5.2Qwen3 8 (max)
BenchLeader Index61.663.6
Agents & tools score61.256.0
Coding score53.769.7
Composite score67.570.9
Human preference score67.968.3
Instruction following score67.4
Knowledge score61.162.3
Long context score67.965.7
Multimodal score59.967.2
Reasoning score61.368.1
Maths score62.4
Blended price $/M$4.81$2.67
Output speed68 tok/s38 tok/s
Time to first answer129.4 s55.3 s
Context window400k1M
Humanity's Last Exam27.8%
Terminal-Bench64.9%
SimpleBench45.8%
SciCode52.9%
Remote Labor Index2.1%
APEX-Agents23.0%
ProofBench58.0%
Epoch Capabilities Index153.5156.6
LMArena Text14761480
LMArena Hard Prompts14971502
LMArena Coding15151520
LMArena WebDev14171670
LMArena Vision12681313
LMArena Agent3.9
LiveBench78.5%
LiveBench Reasoning88.2%
LiveBench Coding72.9%
LiveBench Agentic Coding64.7%
LiveBench Mathematics91.3%
LiveBench Data Analysis78.4%
LiveBench Language79.7%
AA Intelligence Index30.440.3
IFBench75.4%
AA-LCR82.7%78.3%
MMMU-Pro82.3%
AA-Omniscience-0.93.4
Terminal-Bench Hard47.0%
GPQA Diamond (AA)90.3%92.7%
Humanity's Last Exam (AA)37.7%43.0%
SciCode (AA)53.2%
τ²-Bench Telecom (AA)84.8%
LiveCodeBench87.8%
MMLU-Pro88.6%
IOI73.0%
LegalBench83.6%
CorpFin65.8%
TaxEval75.5%
Terminal-Bench 2.1 (Vals)67.4%
SWE-bench (Vals)85.6%
GPQA Diamond (Vals)93.7%
Vals Index51.8
SWE-Bench Pro29.9%
VISTA46.6%
MultiNRC42.2%
TutorBench53.5%
Kagi LLM Benchmark73.3%
SWE-bench Verified (bash only)69.0%
ARC-AGI-194.5%
ARC-AGI-272.9%
BFCL Overall55.9%

Data as of 2026-09-10. Best configuration of each model; every score links to its source on the model pages.

GPT-5.2 vs Qwen3 8: questions

Is GPT-5.2 better than Qwen3 8?
Qwen3 8 (max) leads on quality: 63.6 vs 61.6. The BenchLeader Index combines every independent quality benchmark; Qwen3 8 (max) is ahead overall as of 2026-09-10, but check the category scores for your use.
Is GPT-5.2 better than Qwen3 8 for coding?
Qwen3 8 scores higher in coding (70 vs 54 on the category index, where 50 is average).
Is GPT-5.2 better than Qwen3 8 for agentic tasks?
GPT-5.2 scores higher in agentic tasks (61 vs 56 on the category index, where 50 is average).
Which is cheaper, GPT-5.2 or Qwen3 8?
Qwen3 8 is cheaper: $2.67 against $4.81 per million tokens, blended at three input tokens per output token.
Which is faster, GPT-5.2 or Qwen3 8?
GPT-5.2 streams faster: 68 against 38 output tokens per second.
Which has the larger context window?
Qwen3 8 accepts more context: 1M against 400k tokens.