BenchLeader

Claude Haiku 5.5 vs Gemini 4 Argon

Verdict
  • Gemini 4 Argon (high) leads on quality: 70.0 vs 61.3.
  • Gemini 4 Argon (high) is stronger in coding, composite, knowledge, long context, maths, reasoning, agents & tools, human preference.
  • Claude Haiku 5.5 (high) is 20× cheaper ($0.200 vs $4.00 per 1M blended).
MetricClaude Haiku 5.5 (high)Gemini 4 Argon (high)
BenchLeader Index61.370.0
Coding score62.672.1
Composite score72.489.2
Knowledge score64.375.7
Long context score63.865.0
Maths score66.271.9
Reasoning score72.182.0
Agents & tools score–67.5
Human preference score–73.2
Blended price $/M$0.200$4.00
Output speed174 tok/s–
Time to first answer22.6 s–
Context window1M1M
OTIS Mock AIME97.2%–
SciCode–61.8%
LMArena Text–1525
LMArena Hard Prompts–1551
LMArena Coding–1562
LMArena WebDev15871678
LMArena Agent–9.3
AA Intelligence Index v4.3.237.852.6
AA-LCR77.3%79.7%
AA-Omniscience5.842.4
Humanity's Last Exam (AA)37.3%57.1%
SciCode (AA)48.7%61.8%
IOI–100.0%
LegalBench–88.3%
Vals Index–68.9
CritPt18.6%27.1%
GDPval-AA v2.145.9%56.3%
BioMysteryBench–76.3%
Code Migration–68.2%
CUA-bench–4.8%
CyberBench–77.9%
Excel Modeling Benchmark–75.2%
Finance Agent v2–65.4%
Harvey's Legal Agent Benchmark–19.6%
Legal Research Bench–54.8%
MedCode–58.8%
MedScribe–87.4%
MysteryMechanism–45.5%
ProgramBench–2.5%
Public Benefits Bench–69.8%
SAGE–53.6%
SREBench–44.3%
Tax Agent Bench–76.2%
Terminal-Bench 4.0 (Vals)–57.6%
Terminal-Bench Science–44.3%
Vibe Code Bench v1.1–91.9%
LMArena Maths–1528
LMArena Creative Writing–1519
LMArena Instruction Following–1529
LMArena Multi-turn–1553
LMArena Longer Queries–1545
FrontierSWE–55.0%
Terminal-Bench 4.0 (AA)21.7%57.1%
AutomationBench33.7%77.5%
GDP.pdf17.2%21.8%
Harvey LAB1.1%–
AA-Omniscience: accuracy34.8%49.9%
AA-Omniscience: non-hallucination55.5%84.9%
AA-Briefcase v1.114421488

Data as of 2026-10-11. Best configuration of each model; every score links to its source on the model pages.

Claude Haiku 5.5 vs Gemini 4 Argon: questions

Is Claude Haiku 5.5 better than Gemini 4 Argon?
Gemini 4 Argon (high) leads on quality: 70.0 vs 61.3. The BenchLeader Index combines every independent quality benchmark; Gemini 4 Argon (high) is ahead overall as of 2026-10-11, but check the category scores for your use.
Is Claude Haiku 5.5 better than Gemini 4 Argon for coding?
Gemini 4 Argon scores higher in coding (72 vs 63 on the category index, where 50 is average).
Which is cheaper, Claude Haiku 5.5 or Gemini 4 Argon?
Claude Haiku 5.5 is cheaper: $0.200 against $4.00 per million tokens, blended at three input tokens per output token.
Which has the larger context window?
Both accept 1M tokens of context.