BenchLeader

GPT-5.6 Sol vs Qwen3.8 2.4T A95B

Verdict
  • GPT-5.6 Sol (max) leads on quality: 68.8 vs 65.2.
  • GPT-5.6 Sol (max) is stronger in coding, instruction following, knowledge, long context, maths, multimodal.
  • Qwen3.8 2.4T A95B is stronger in agents & tools, composite, reasoning.
  • Qwen3.8 2.4T A95B is 2.7× cheaper ($3.00 vs $8.00 per 1M blended).
  • GPT-5.6 Sol (max) streams 1.6× faster (61 vs 38 tokens per second).
MetricGPT-5.6 Sol (max)Qwen3.8 2.4T A95B
BenchLeader Index68.865.2
Agents & tools score66.678.3
Coding score67.0
Composite score79.279.8
Instruction following score71.0
Knowledge score66.566.3
Long context score68.566.6
Maths score74.3
Multimodal score68.1
Reasoning score73.780.5
Blended price $/M$8.00$3.00
Output speed61 tok/s38 tok/s
Time to first answer127.3 s55.2 s
Context window1.1M984k
GPQA Diamond93.5%
FrontierMath Tiers 1–389.1%
FrontierMath Tier 482.9%
OTIS Mock AIME100.0%
SimpleQA Verified69.7%
Terminal-Bench37.3%
OSWorld-Verified 2.027.3%
SciCode56.1%
WeirdML87.0%
APEX-Agents39.9%
ProofBench83.0%
LiveBench81.0%
LiveBench Reasoning91.7%
LiveBench Coding83.9%
LiveBench Agentic Coding56.2%
LiveBench Mathematics96.2%
LiveBench Data Analysis79.8%
LiveBench Language87.7%
LiveBench Instruction Following71.8%
AA Intelligence Index47.140.0
IFBench72.7%
AA-LCR84.0%80.3%
MMMU-Pro83.4%
AA-Omniscience22.04.3
Terminal-Bench Hard65.9%
GPQA Diamond (AA)94.1%93.5%
Humanity's Last Exam (AA)49.5%42.5%
SciCode (AA)57.1%54.0%
τ²-Bench Telecom (AA)85.1%
LiveCodeBench82.6%
MMLU-Pro89.1%
IOI91.2%
LegalBench87.0%
CorpFin64.4%
TaxEval74.8%
Terminal-Bench 2.1 (Vals)85.8%
SWE-bench (Vals)96.2%
GPQA Diamond (Vals)95.2%
Vals Index63.7
PRBench Finance50.5%
PRBench Legal50.5%
ARC-AGI-196.5%
ARC-AGI-292.5%
ARC-AGI-37.8%
CritPt32.3%20.0%
GDPval (AA)54.3%56.4%
τ²-Bench Banking (AA)44.3%49.1%
ITBench SRE (AA)56.2%
Analyst Agent (AA)47.5%
BioMysteryBench71.1%
Code Migration52.9%
CUA-bench8.3%
Excel Modeling Benchmark72.3%
Finance Agent v253.8%
Harvey's Legal Agent Benchmark2.5%
Legal Research Bench48.1%
MedCode44.0%
MedScribe85.2%
MMMU-Pro (Vals)88.8%
MortgageTax67.3%
MysteryMechanism33.3%
ProgramBench1.5%
Public Benefits Bench66.5%
SAGE52.6%
SkillsBench54.1%
SREBench30.5%
Tax Agent Bench68.0%
Terminal-Bench 4.0 (Vals)27.8%
Terminal-Bench Science12.9%
Time Horizon Index: KSP23.8%
Vals Multimodal Index72.6%
Vibe Code Bench 1-10020.0%
Vibe Code Bench v1.180.5%
Web Search Index43.6%
DrugDiscoveryBench56.5%
Chess Puzzles55.0%
EBR-bench44.8%
Mystery Game Puzzles58.0%
PostTrainBench36.2%
DeepSWE72.7%
LMCA58.4%
DTBench95.5%
CursorBench67.2%
ALE-Bench2176.9
GDP.pdf30.7%
FrontierSWE32.2%
BTF-313.7%

Data as of 2026-09-19. Best configuration of each model; every score links to its source on the model pages.

GPT-5.6 Sol vs Qwen3.8 2.4T A95B: questions

Is GPT-5.6 Sol better than Qwen3.8 2.4T A95B?
GPT-5.6 Sol (max) leads on quality: 68.8 vs 65.2. The BenchLeader Index combines every independent quality benchmark; GPT-5.6 Sol (max) is ahead overall as of 2026-09-19, but check the category scores for your use.
Is GPT-5.6 Sol better than Qwen3.8 2.4T A95B for agentic tasks?
Qwen3.8 2.4T A95B scores higher in agentic tasks (78 vs 67 on the category index, where 50 is average).
Which is cheaper, GPT-5.6 Sol or Qwen3.8 2.4T A95B?
Qwen3.8 2.4T A95B is cheaper: $3.00 against $8.00 per million tokens, blended at three input tokens per output token.
Which is faster, GPT-5.6 Sol or Qwen3.8 2.4T A95B?
GPT-5.6 Sol streams faster: 61 against 38 output tokens per second.
Which has the larger context window?
GPT-5.6 Sol accepts more context: 1.1M against 984k tokens.