BenchLeader

Claude Haiku 5.5 vs GPT-6 Astra

Verdict
  • GPT-6 Astra (max) leads on quality: 70.2 vs 61.3.
  • GPT-6 Astra (max) is stronger in coding, composite, knowledge, long context, maths, reasoning, agents & tools, human preference, multimodal.
  • Claude Haiku 5.5 (high) is 100× cheaper ($0.200 vs $20.00 per 1M blended).
  • Claude Haiku 5.5 (high) streams 3.7× faster (174 vs 47 tokens per second).
MetricClaude Haiku 5.5 (high)GPT-6 Astra (max)
BenchLeader Index61.370.2
Coding score62.674.4
Composite score72.480.3
Knowledge score64.379.9
Long context score63.865.6
Maths score66.272.2
Reasoning score72.176.7
Agents & tools score–68.8
Human preference score–67.0
Multimodal score–66.3
Blended price $/M$0.200$20.00
Output speed174 tok/s47 tok/s
Time to first answer22.6 s383.6 s
Context window1M1.1M
GPQA Diamond–95.8%
FrontierMath Tiers 1–3–93.7%
FrontierMath Tier 4–97.6%
OTIS Mock AIME97.2%100.0%
SimpleQA Verified–75.6%
Terminal-Bench–58.2%
SciCode–56.5%
WeirdML–93.3%
FrontierCode–53.3%
LMArena Text–1475
LMArena Hard Prompts–1500
LMArena Coding–1542
LMArena WebDev15871786
LMArena Vision–1281
LMArena Agent–13.1
LiveBench–82.2%
LiveBench Reasoning–92.7%
LiveBench Coding–80.4%
LiveBench Agentic Coding–57.3%
LiveBench Mathematics–96.8%
LiveBench Data Analysis–83.0%
LiveBench Language–89.4%
LiveBench Instruction Following–75.6%
AA Intelligence Index v4.3.237.852.7
AA-LCR77.3%80.7%
MMMU-Pro–86.9%
AA-Omniscience5.843.4
GPQA Diamond (AA)–96.1%
Humanity's Last Exam (AA)37.3%54.7%
SciCode (AA)48.7%56.5%
IOI–100.0%
Terminal-Bench 2.1 (Vals)–87.3%
Vals Index–63.1
ARC-AGI-1–97.5%
ARC-AGI-2–95.0%
ARC-AGI-3–98.6%
CritPt18.6%31.7%
GDPval-AA v2.145.9%52.1%
τ³-Banking (AA)–41.4%
ITBench SRE (AA)–48.6%
Analyst Agent (AA)–51.3%
BioMysteryBench–79.3%
Code Migration–67.7%
CUA-bench–19.2%
CyberBench–41.1%
Excel Modeling Benchmark–71.7%
Finance Agent v2–53.5%
Harvey's Legal Agent Benchmark–5.4%
Legal Research Bench–39.4%
MedCode–48.5%
MedScribe–87.9%
MysteryMechanism–53.1%
ProgramBench–5.5%
SAGE–46.4%
SREBench–56.9%
Tax Agent Bench–63.3%
Terminal-Bench 4.0 (Vals)–59.6%
Terminal-Bench Science–62.9%
Time Horizon Index: KSP–90.5%
Vibe Code Bench 1-100–27.6%
Vibe Code Bench v1.1–89.6%
LMArena Maths–1486
LMArena Creative Writing–1448
LMArena Instruction Following–1469
LMArena Multi-turn–1484
LMArena Longer Queries–1484
LMArena Document–1473
Chess Puzzles–72.0%
EBR-bench–76.2%
Mystery Game Puzzles–84.0%
BALROG–68.3%
DeepSWE v1.1–73.2%
ALE-Bench–2951.3
GDP.pdf–34.2%
FrontierSWE–65.5%
Terminal-Bench 4.0 (AA)21.7%59.1%
Terminal-Bench 2.1 (AA)–88.4%
AutomationBench33.7%68.5%
GDP.pdf17.2%31.0%
MLCR–35.0%
Harvey LAB1.1%8.6%
AA-Omniscience: accuracy34.8%62.6%
AA-Omniscience: non-hallucination55.5%48.7%
AA-Briefcase v1.114421570

Data as of 2026-10-11. Best configuration of each model; every score links to its source on the model pages.

Claude Haiku 5.5 vs GPT-6 Astra: questions

Is Claude Haiku 5.5 better than GPT-6 Astra?
GPT-6 Astra (max) leads on quality: 70.2 vs 61.3. The BenchLeader Index combines every independent quality benchmark; GPT-6 Astra (max) is ahead overall as of 2026-10-11, but check the category scores for your use.
Is Claude Haiku 5.5 better than GPT-6 Astra for coding?
GPT-6 Astra scores higher in coding (74 vs 63 on the category index, where 50 is average).
Which is cheaper, Claude Haiku 5.5 or GPT-6 Astra?
Claude Haiku 5.5 is cheaper: $0.200 against $20.00 per million tokens, blended at three input tokens per output token.
Which is faster, Claude Haiku 5.5 or GPT-6 Astra?
Claude Haiku 5.5 streams faster: 174 against 47 output tokens per second.
Which has the larger context window?
GPT-6 Astra accepts more context: 1.1M against 1M tokens.