BenchLeader

Claude Fable 5 vs Grok 4.20 Multi-Agent

Verdict
  • Claude Fable 5 (thinking) leads on quality: 70.8 vs 59.2.
  • Claude Fable 5 (thinking) is stronger in agents & tools, composite, instruction following, knowledge, long context, reasoning.
  • Grok 4.20 Multi-Agent is stronger in coding, human preference, maths, multimodal.
  • Grok 4.20 Multi-Agent is 13× cheaper ($1.56 vs $20.00 per 1M blended).
  • Grok 4.20 Multi-Agent streams 4.4× faster (304 vs 70 tokens per second).
MetricClaude Fable 5 (thinking)Grok 4.20 Multi-Agent
BenchLeader Index70.859.2
Agents & tools score74.1
Composite score91.7
Instruction following score63.0
Knowledge score84.6
Long context score67.6
Reasoning score90.865.9
Coding score65.5
Human preference score67.0
Maths score64.4
Multimodal score60.4
Blended price $/M$20.00$1.56
Output speed70 tok/s304 tok/s
Time to first answer135.7 s9.9 s
Context window1M1M
LMArena Text1470
LMArena Hard Prompts1483
LMArena Coding1508
LMArena Vision1260
AA Intelligence Index49.6
IFBench63.5%
AA-LCR82.3%
AA-Omniscience43.3
Terminal-Bench Hard62.9%
GPQA Diamond (AA)92.6%
Humanity's Last Exam (AA)55.5%
SciCode (AA)61.0%
τ²-Bench Telecom (AA)98.5%
Kagi LLM Benchmark91.4%
CritPt28.6%
GDPval (AA)54.8%
τ²-Bench Banking (AA)38.1%
Analyst Agent (AA)48.8%
LMArena Maths1453
LMArena Creative Writing1448
LMArena Instruction Following1443
LMArena Multi-turn1473
LMArena Longer Queries1457
LMArena Search1205

Data as of 2026-09-21. Best configuration of each model; every score links to its source on the model pages.

Claude Fable 5 vs Grok 4.20 Multi-Agent: questions

Is Claude Fable 5 better than Grok 4.20 Multi-Agent?
Claude Fable 5 (thinking) leads on quality: 70.8 vs 59.2. The BenchLeader Index combines every independent quality benchmark; Claude Fable 5 (thinking) is ahead overall as of 2026-09-21, but check the category scores for your use.
Which is cheaper, Claude Fable 5 or Grok 4.20 Multi-Agent?
Grok 4.20 Multi-Agent is cheaper: $1.56 against $20.00 per million tokens, blended at three input tokens per output token.
Which is faster, Claude Fable 5 or Grok 4.20 Multi-Agent?
Grok 4.20 Multi-Agent streams faster: 304 against 70 output tokens per second.
Which has the larger context window?
Both accept 1M tokens of context.