BenchLeader

GPT-6 Astra vs Qwen3.5 397B A17B

Verdict
  • GPT-6 Astra (max) leads on quality: 70.9 vs 58.7.
  • GPT-6 Astra (max) is stronger in coding, composite, knowledge, maths, reasoning.
  • Qwen3.5 397B A17B is stronger in agents & tools, human preference, instruction following, long context, multimodal.
  • Qwen3.5 397B A17B is 52× cheaper ($0.387 vs $20.00 per 1M blended).
  • Qwen3.5 397B A17B streams 2.9× faster (78 vs 27 tokens per second).
MetricGPT-6 Astra (max)Qwen3.5 397B A17B
BenchLeader Index70.958.7
Agents & tools score68.269.4
Coding score78.651.7
Composite score72.953.1
Knowledge score80.049.5
Maths score77.250.1
Reasoning score73.663.0
Human preference score63.6
Instruction following score76.1
Long context score65.2
Multimodal score61.6
Blended price $/M$20.00$0.387
Output speed27 tok/s78 tok/s
Time to first answer337.3 s42.7 s
Context window1.1M262k
GPQA Diamond95.8%85.9%
FrontierMath Tiers 1–393.7%29.5%
FrontierMath Tier 497.6%
OTIS Mock AIME100.0%88.9%
SimpleQA Verified75.6%
Terminal-Bench58.2%
SciCode56.5%
FrontierCode53.3%
Epoch Capabilities Index147.0
LMArena Text1441
LMArena Hard Prompts1463
LMArena Coding1491
LMArena WebDev17961399
LMArena Vision1265
LMArena Agent12.5
LiveBench82.2%
LiveBench Reasoning92.7%
LiveBench Coding80.4%
LiveBench Agentic Coding57.3%
LiveBench Mathematics96.8%
LiveBench Data Analysis83.0%
LiveBench Language89.4%
AA Intelligence Index19.1
IFBench78.8%
AA-LCR77.3%
MMMU-Pro77.3%
AA-Omniscience-30.8
Terminal-Bench Hard40.9%
GPQA Diamond (AA)89.3%
Humanity's Last Exam (AA)29.0%
SciCode (AA)44.8%
τ²-Bench Telecom (AA)95.6%
Terminal-Bench 2.1 (Vals)87.3%
Vals Index66.6
AIME 202694.2%
HMMT February 202687.9%
ARC-AGI-197.5%
ARC-AGI-295.0%
ARC-AGI-398.6%

Data as of 2026-09-10. Best configuration of each model; every score links to its source on the model pages.

GPT-6 Astra vs Qwen3.5 397B A17B: questions

Is GPT-6 Astra better than Qwen3.5 397B A17B?
GPT-6 Astra (max) leads on quality: 70.9 vs 58.7. The BenchLeader Index combines every independent quality benchmark; GPT-6 Astra (max) is ahead overall as of 2026-09-10, but check the category scores for your use.
Is GPT-6 Astra better than Qwen3.5 397B A17B for coding?
GPT-6 Astra scores higher in coding (79 vs 52 on the category index, where 50 is average).
Is GPT-6 Astra better than Qwen3.5 397B A17B for agentic tasks?
Qwen3.5 397B A17B scores higher in agentic tasks (69 vs 68 on the category index, where 50 is average).
Which is cheaper, GPT-6 Astra or Qwen3.5 397B A17B?
Qwen3.5 397B A17B is cheaper: $0.387 against $20.00 per million tokens, blended at three input tokens per output token.
Which is faster, GPT-6 Astra or Qwen3.5 397B A17B?
Qwen3.5 397B A17B streams faster: 78 against 27 output tokens per second.
Which has the larger context window?
GPT-6 Astra accepts more context: 1.1M against 262k tokens.