BenchLeader

GPT-5.4 vs Qwen3.8 2.4T A95B

Verdict
  • GPT-5.4 (xhigh) and Qwen3.8 2.4T A95B are level on quality (65.4 vs 65.2).
  • GPT-5.4 (xhigh) is stronger in coding, instruction following, long context, maths, multimodal.
  • Qwen3.8 2.4T A95B is stronger in agents & tools, composite, knowledge, reasoning.
  • Qwen3.8 2.4T A95B is 1.9× cheaper ($3.00 vs $5.63 per 1M blended).
  • GPT-5.4 (xhigh) streams 3.2× faster (121 vs 38 tokens per second).
MetricGPT-5.4 (xhigh)Qwen3.8 2.4T A95B
BenchLeader Index65.465.2
Agents & tools score63.478.3
Coding score66.6
Composite score69.479.8
Instruction following score72.1
Knowledge score60.666.3
Long context score67.566.6
Maths score66.4
Multimodal score62.9
Reasoning score72.480.5
Blended price $/M$5.63$3.00
Output speed121 tok/s38 tok/s
Time to first answer141.3 s55.2 s
Context window1.1M984k
GPQA Diamond93.3%
FrontierMath Tiers 1–378.6%
FrontierMath Tier 449.0%
OTIS Mock AIME95.3%
SimpleQA Verified45.1%
Humanity's Last Exam36.2%
SciCode56.6%
WeirdML77.7%
APEX-Agents36.0%
ProofBench56.0%
GSO-Bench31.4%
LiveBench78.0%
LiveBench Reasoning88.1%
LiveBench Coding77.5%
LiveBench Agentic Coding53.8%
LiveBench Mathematics94.2%
LiveBench Data Analysis79.3%
LiveBench Language82.6%
LiveBench Instruction Following70.2%
AA Intelligence Index39.040.0
IFBench74.0%
AA-LCR82.0%80.3%
MMMU-Pro78.4%
AA-Omniscience5.84.3
Terminal-Bench Hard57.6%
GPQA Diamond (AA)92.0%93.5%
Humanity's Last Exam (AA)43.7%42.5%
SciCode (AA)54.0%
τ²-Bench Telecom (AA)87.1%
AIME (Vals)96.7%
LiveCodeBench84.1%
MMLU-Pro87.5%
LegalBench86.0%
CorpFin65.3%
TaxEval74.0%
MedQA96.1%
SWE-bench (Vals)78.2%
GPQA Diamond (Vals)91.7%
AIME 202699.2%
HMMT February 202697.7%
MathArena Apex54.2%
SWE-Bench Pro59.1%
MCP Atlas70.6%
ARC-AGI-193.7%
ARC-AGI-274.0%
CritPt23.4%20.0%
GDPval (AA)40.3%56.4%
τ²-Bench Banking (AA)39.6%49.1%
APEX-Agents (AA)33.3%
CaseLaw v263.8%
Code Migration35.0%
Harvey's Legal Agent Benchmark0.0%
MedCode41.3%
MedScribe77.5%
MMMU-Pro (Vals)87.5%
MortgageTax68.3%
ProgramBench0.0%
SAGE43.3%
SkillsBench51.7%
Vibe Code Bench v1.167.4%
SWE-Bench Pro (private)43.4%
SWE Atlas: Codebase QnA40.8%
SWE Atlas: Refactoring44.3%
SWE Atlas: Test Writing44.4%
Chess Puzzles44.0%
EBR-bench25.4%
Mystery Game Puzzles37.0%
CL-bench27.9%
CL-bench Life21.7%
METR Time Horizons74.3%
DeepSWE51.8%
LMCA52.0%
DTBench94.4%

Data as of 2026-09-19. Best configuration of each model; every score links to its source on the model pages.

GPT-5.4 vs Qwen3.8 2.4T A95B: questions

Is GPT-5.4 better than Qwen3.8 2.4T A95B?
GPT-5.4 (xhigh) and Qwen3.8 2.4T A95B are level on quality (65.4 vs 65.2). The BenchLeader Index combines every independent quality benchmark; GPT-5.4 (xhigh) is ahead overall as of 2026-09-19, but check the category scores for your use.
Is GPT-5.4 better than Qwen3.8 2.4T A95B for agentic tasks?
Qwen3.8 2.4T A95B scores higher in agentic tasks (78 vs 63 on the category index, where 50 is average).
Which is cheaper, GPT-5.4 or Qwen3.8 2.4T A95B?
Qwen3.8 2.4T A95B is cheaper: $3.00 against $5.63 per million tokens, blended at three input tokens per output token.
Which is faster, GPT-5.4 or Qwen3.8 2.4T A95B?
GPT-5.4 streams faster: 121 against 38 output tokens per second.
Which has the larger context window?
GPT-5.4 accepts more context: 1.1M against 984k tokens.