BenchLeader

Claude Fable 5.1 vs Grok 4.20

Verdict
  • Claude Fable 5.1 (high) leads on quality: 72.0 vs 59.9.
  • Claude Fable 5.1 (high) is stronger in agents & tools, coding, composite, knowledge, long context, reasoning.
  • Grok 4.20 (thinking) is stronger in human preference, instruction following, maths, multimodal.
  • Grok 4.20 (thinking) is 13× cheaper ($1.56 vs $20.00 per 1M blended).
  • Grok 4.20 (thinking) streams 1.8× faster (101 vs 55 tokens per second).
MetricClaude Fable 5.1 (high)Grok 4.20 (thinking)
BenchLeader Index72.059.9
Agents & tools score66.651.7
Coding score74.154.7
Composite score94.061.5
Knowledge score83.756.8
Long context score68.460.7
Reasoning score81.060.7
Human preference score67.1
Instruction following score79.9
Maths score55.6
Multimodal score59.9
Blended price $/M$20.00$1.56
Output speed55 tok/s101 tok/s
Time to first answer14.7 s21.4 s
Context window1M1M
GPQA Diamond89.3%
FrontierMath Tiers 1–344.9%
FrontierMath Tier 417.1%
OTIS Mock AIME92.2%
SimpleQA Verified30.2%
Terminal-Bench54.5%
SciCode57.6%
WeirdML92.3%
APEX-Agents44.4%
ProofBench14.0%
LMArena Text1472
LMArena Hard Prompts1488
LMArena Coding1510
LMArena WebDev1374
LMArena Vision1263
AA Intelligence Index51.225.7
IFBench82.9%
AA-LCR83.7%69.0%
MMMU-Pro74.6%
AA-Omniscience40.814.8
Terminal-Bench Hard40.9%
GPQA Diamond (AA)90.6%91.1%
Humanity's Last Exam (AA)55.9%34.5%
SciCode (AA)58.7%
τ²-Bench Telecom (AA)96.5%
AIME (Vals)96.5%
LiveCodeBench84.3%
MMLU-Pro86.3%
LegalBench77.7%
CorpFin63.7%
TaxEval74.1%
MedQA94.5%
Terminal-Bench 2.1 (Vals)44.2%
SWE-bench (Vals)72.2%
GPQA Diamond (Vals)88.6%
Vals Index17.6
Kagi LLM Benchmark75.0%
ARC-AGI-196.0%89.5%
ARC-AGI-288.8%65.1%
ARC-AGI-30.1%
CritPt30.3%6.6%
GDPval (AA)57.5%
τ²-Bench Banking (AA)43.1%
APEX-Agents (AA)14.2%
CaseLaw v254.5%
Code Migration0.3%
Excel Modeling Benchmark11.7%
Finance Agent v228.5%
Harvey's Legal Agent Benchmark0.0%
Legal Research Bench13.9%
MedCode32.2%
MedScribe63.4%
MMMU-Pro (Vals)83.5%
MortgageTax45.4%
SAGE38.2%
Terminal-Bench 2.0 (Vals)40.5%
Vals Multimodal Index39.1%
Vibe Code Bench v1.14.1%
LMArena Maths1467
LMArena Creative Writing1446
LMArena Instruction Following1445
LMArena Multi-turn1480
LMArena Longer Queries1465
LMArena Document1439
Chess Puzzles24.0%
MirrorCode73.3%
ForecastBench60.7%
CursorBench69.4%
ALE-Bench2143.2

Data as of 2026-09-19. Best configuration of each model; every score links to its source on the model pages.

Claude Fable 5.1 vs Grok 4.20: questions

Is Claude Fable 5.1 better than Grok 4.20?
Claude Fable 5.1 (high) leads on quality: 72.0 vs 59.9. The BenchLeader Index combines every independent quality benchmark; Claude Fable 5.1 (high) is ahead overall as of 2026-09-19, but check the category scores for your use.
Is Claude Fable 5.1 better than Grok 4.20 for coding?
Claude Fable 5.1 scores higher in coding (74 vs 55 on the category index, where 50 is average).
Is Claude Fable 5.1 better than Grok 4.20 for agentic tasks?
Claude Fable 5.1 scores higher in agentic tasks (67 vs 52 on the category index, where 50 is average).
Which is cheaper, Claude Fable 5.1 or Grok 4.20?
Grok 4.20 is cheaper: $1.56 against $20.00 per million tokens, blended at three input tokens per output token.
Which is faster, Claude Fable 5.1 or Grok 4.20?
Grok 4.20 streams faster: 101 against 55 output tokens per second.
Which has the larger context window?
Both accept 1M tokens of context.