BenchLeader

GPT-6.1 Sol vs Qwen3.8 2.4T A95B

Verdict
  • GPT-6.1 Sol (max) leads on quality: 69.3 vs 63.3.
  • GPT-6.1 Sol (max) is stronger in coding, composite, human preference, knowledge, long context, maths, multimodal, reasoning.
  • Qwen3.8 2.4T A95B is stronger in agents & tools.
  • Qwen3.8 2.4T A95B is 1.3× cheaper ($3.00 vs $4.00 per 1M blended).
  • GPT-6.1 Sol (max) streams 1.5× faster (56 vs 37 tokens per second).
MetricGPT-6.1 Sol (max)Qwen3.8 2.4T A95B
BenchLeader Index69.363.3
Agents & tools score63.978.5
Coding score72.4–
Composite score79.074.7
Human preference score68.1–
Knowledge score78.863.6
Long context score66.865.4
Maths score72.5–
Multimodal score66.2–
Reasoning score76.474.5
Blended price $/M$4.00$3.00
Output speed56 tok/s37 tok/s
Time to first answer326.9 s57.8 s
Context window1.1M1M
GPQA Diamond95.4%–
FrontierMath Tiers 1–393.7%–
FrontierMath Tier 4100.0%–
OTIS Mock AIME100.0%–
SimpleQA Verified73.9%–
Terminal-Bench58.2%–
SciCode54.2%–
APEX-Agents60.0%–
LMArena Text1484–
LMArena Hard Prompts1507–
LMArena Coding1545–
LMArena WebDev1755–
LMArena Vision1288–
LMArena Agent11.7–
LiveBench81.6%–
LiveBench Reasoning92.6%–
LiveBench Coding80.4%–
LiveBench Agentic Coding54.5%–
LiveBench Mathematics96.8%–
LiveBench Data Analysis82.7%–
LiveBench Language90.1%–
LiveBench Instruction Following74.2%–
AA Intelligence Index v4.3.251.839.9
AA-LCR83.0%80.3%
MMMU-Pro86.0%–
AA-Omniscience41.54.3
GPQA Diamond (AA)–93.5%
Humanity's Last Exam (AA)52.9%42.5%
SciCode (AA)54.2%54.0%
IOI96.9%–
Vals Index61.1–
ARC-AGI-196.5%–
ARC-AGI-294.2%–
ARC-AGI-396.2%–
CritPt31.7%20.0%
GDPval-AA v2.153.8%55.6%
τ³-Banking (AA)–49.1%
Analyst Agent (AA)50.0%–
BioMysteryBench79.6%–
Code Migration65.1%–
CyberBench39.3%–
Excel Modeling Benchmark70.8%–
Finance Agent v252.0%–
Harvey's Legal Agent Benchmark5.4%–
Legal Research Bench38.5%–
MedCode48.8%–
MedScribe86.5%–
MysteryMechanism46.4%–
Public Benefits Bench59.3%–
SAGE46.5%–
SREBench50.8%–
Tax Agent Bench62.3%–
Terminal-Bench 4.0 (Vals)55.0%–
Vibe Code Bench v1.188.9%–
LMArena Maths1488–
LMArena Creative Writing1462–
LMArena Instruction Following1488–
LMArena Multi-turn1487–
LMArena Longer Queries1496–
Chess Puzzles61.0%–
EBR-bench54.3%–
Mystery Game Puzzles80.0%–
Terminal-Bench 4.0 (AA)56.1%11.1%
Terminal-Bench 2.1 (AA)–82.0%
AutomationBench64.9%–
GDP.pdf31.0%–
MLCR33.9%–
Harvey LAB6.9%–
AA-Omniscience: accuracy62.1%31.3%
AA-Omniscience: non-hallucination45.7%60.8%
AA-Briefcase v1.11557–

Data as of 2026-10-11. Best configuration of each model; every score links to its source on the model pages.

GPT-6.1 Sol vs Qwen3.8 2.4T A95B: questions

Is GPT-6.1 Sol better than Qwen3.8 2.4T A95B?
GPT-6.1 Sol (max) leads on quality: 69.3 vs 63.3. The BenchLeader Index combines every independent quality benchmark; GPT-6.1 Sol (max) is ahead overall as of 2026-10-11, but check the category scores for your use.
Is GPT-6.1 Sol better than Qwen3.8 2.4T A95B for agentic tasks?
Qwen3.8 2.4T A95B scores higher in agentic tasks (79 vs 64 on the category index, where 50 is average).
Which is cheaper, GPT-6.1 Sol or Qwen3.8 2.4T A95B?
Qwen3.8 2.4T A95B is cheaper: $3.00 against $4.00 per million tokens, blended at three input tokens per output token.
Which is faster, GPT-6.1 Sol or Qwen3.8 2.4T A95B?
GPT-6.1 Sol streams faster: 56 against 37 output tokens per second.
Which has the larger context window?
GPT-6.1 Sol accepts more context: 1.1M against 1M tokens.