BenchLeader

Claude Fable 5 vs Gemini 4 Argon

Verdict
  • Claude Fable 5 (max) and Gemini 4 Argon (high) are level on quality (69.1 vs 70.0).
  • Claude Fable 5 (max) is stronger in coding, instruction following, knowledge, long context, maths.
  • Gemini 4 Argon (high) is stronger in agents & tools, composite, reasoning, human preference.
  • Gemini 4 Argon (high) is 5.0× cheaper ($4.00 vs $20.00 per 1M blended).
MetricClaude Fable 5 (max)Gemini 4 Argon (high)
BenchLeader Index69.170.0
Agents & tools score65.067.5
Coding score74.172.1
Composite score79.889.2
Instruction following score63.0–
Knowledge score81.275.7
Long context score66.465.0
Maths score73.071.9
Reasoning score74.082.0
Human preference score–73.2
Blended price $/M$20.00$4.00
Output speed65 tok/s–
Time to first answer117.7 s–
Context window1M1M
GPQA Diamond85.9%–
FrontierMath Tiers 1–387.0%–
FrontierMath Tier 490.2%–
OTIS Mock AIME99.7%–
Terminal-Bench44.5%–
SimpleBench81.9%–
SciCode61.0%61.8%
WeirdML91.9%–
ProofBench95.0%–
LMArena Text–1525
LMArena Hard Prompts–1551
LMArena Coding–1562
LMArena WebDev–1678
LMArena Agent–9.3
LiveBench83.0%–
LiveBench Reasoning89.7%–
LiveBench Coding86.0%–
LiveBench Agentic Coding62.2%–
LiveBench Mathematics96.0%–
LiveBench Data Analysis80.5%–
LiveBench Language90.7%–
LiveBench Instruction Following75.8%–
AA Intelligence Index v4.3.249.652.6
IFBench63.5%–
AA-LCR82.3%79.7%
AA-Omniscience43.342.4
Terminal-Bench Hard62.9%–
GPQA Diamond (AA)92.6%–
Humanity's Last Exam (AA)55.5%57.1%
SciCode (AA)61.0%61.8%
τ²-Bench Telecom (AA)98.5%–
IOI–100.0%
LegalBench–88.3%
Vals Index–68.9
ARC-AGI-198.5%–
ARC-AGI-289.2%–
CritPt28.6%27.1%
GDPval-AA v2.155.6%56.3%
τ³-Banking (AA)38.1%–
Analyst Agent (AA)48.8%–
BioMysteryBench–76.3%
Code Migration–68.2%
CUA-bench–4.8%
CyberBench–77.9%
Excel Modeling Benchmark–75.2%
Finance Agent v2–65.4%
Harvey's Legal Agent Benchmark–19.6%
Legal Research Bench–54.8%
MedCode–58.8%
MedScribe–87.4%
MysteryMechanism–45.5%
ProgramBench–2.5%
Public Benefits Bench–69.8%
SAGE–53.6%
SREBench–44.3%
Tax Agent Bench–76.2%
Terminal-Bench 4.0 (Vals)–57.6%
Terminal-Bench Science–44.3%
Vibe Code Bench v1.1–91.9%
LMArena Maths–1528
LMArena Creative Writing–1519
LMArena Instruction Following–1529
LMArena Multi-turn–1553
LMArena Longer Queries–1545
Chess Puzzles41.0%–
EBR-bench39.5%–
Mystery Game Puzzles52.0%–
PostTrainBench41.8%–
DeepSWE v1.169.7%–
LMCA60.3%–
DTBench98.4%–
Vending-Bench 24966.6–
GDP.pdf29.8%–
FrontierSWE47.0%55.0%
Terminal-Bench 4.0 (AA)42.4%57.1%
Terminal-Bench 2.1 (AA)84.6%–
AutomationBench54.1%77.5%
GDP.pdf24.0%21.8%
MLCR64.4%–
EnterpriseOps-Gym51.1%–
AA-Omniscience: accuracy65.3%49.9%
AA-Omniscience: non-hallucination36.4%84.9%
AA-Briefcase v1.115401488

Data as of 2026-10-11. Best configuration of each model; every score links to its source on the model pages.

Claude Fable 5 vs Gemini 4 Argon: questions

Is Claude Fable 5 better than Gemini 4 Argon?
Claude Fable 5 (max) and Gemini 4 Argon (high) are level on quality (69.1 vs 70.0). The BenchLeader Index combines every independent quality benchmark; Gemini 4 Argon (high) is ahead overall as of 2026-10-11, but check the category scores for your use.
Is Claude Fable 5 better than Gemini 4 Argon for coding?
Claude Fable 5 scores higher in coding (74 vs 72 on the category index, where 50 is average).
Is Claude Fable 5 better than Gemini 4 Argon for agentic tasks?
Gemini 4 Argon scores higher in agentic tasks (68 vs 65 on the category index, where 50 is average).
Which is cheaper, Claude Fable 5 or Gemini 4 Argon?
Gemini 4 Argon is cheaper: $4.00 against $20.00 per million tokens, blended at three input tokens per output token.
Which has the larger context window?
Both accept 1M tokens of context.