BenchLeader

GPT-6 Astra vs Qwen3.8 Max (0902)

Verdict
  • GPT-6 Astra (max) leads on quality: 70.5 vs 57.1.
  • GPT-6 Astra (max) is stronger in agents & tools, coding, composite, human preference, knowledge, maths, reasoning.
  • Qwen3.8 Max (0902) (xhigh) is 6.7× cheaper ($3.00 vs $20.00 per 1M blended).
  • GPT-6 Astra (max) streams 1.6× faster (60 vs 38 tokens per second).
MetricGPT-6 Astra (max)Qwen3.8 Max (0902) (xhigh)
BenchLeader Index70.557.1
Agents & tools score68.2
Coding score77.3
Composite score72.9
Human preference score67.9
Knowledge score80.055.9
Maths score77.261.2
Reasoning score72.766.9
Blended price $/M$20.00$3.00
Output speed60 tok/s38 tok/s
Time to first answer321.0 s2.0 s
Context window1.1M1M
GPQA Diamond95.8%92.3%
FrontierMath Tiers 1–393.7%65.6%
FrontierMath Tier 497.6%34.1%
OTIS Mock AIME100.0%100.0%
SimpleQA Verified75.6%47.3%
Terminal-Bench58.2%
SciCode56.5%
FrontierCode53.3%
LMArena Text1478
LMArena Hard Prompts1495
LMArena Coding1537
LMArena WebDev1800
LMArena Agent12.4
LiveBench82.2%
LiveBench Reasoning92.7%
LiveBench Coding80.4%
LiveBench Agentic Coding57.3%
LiveBench Mathematics96.8%
LiveBench Data Analysis83.0%
LiveBench Language89.4%
IOI100.0%
Terminal-Bench 2.1 (Vals)87.3%
Vals Index66.6
ARC-AGI-197.5%
ARC-AGI-295.0%
ARC-AGI-398.6%

Data as of 2026-09-13. Best configuration of each model; every score links to its source on the model pages.

GPT-6 Astra vs Qwen3.8 Max (0902): questions

Is GPT-6 Astra better than Qwen3.8 Max (0902)?
GPT-6 Astra (max) leads on quality: 70.5 vs 57.1. The BenchLeader Index combines every independent quality benchmark; GPT-6 Astra (max) is ahead overall as of 2026-09-13, but check the category scores for your use.
Which is cheaper, GPT-6 Astra or Qwen3.8 Max (0902)?
Qwen3.8 Max (0902) is cheaper: $3.00 against $20.00 per million tokens, blended at three input tokens per output token.
Which is faster, GPT-6 Astra or Qwen3.8 Max (0902)?
GPT-6 Astra streams faster: 60 against 38 output tokens per second.
Which has the larger context window?
GPT-6 Astra accepts more context: 1.1M against 1M tokens.