BenchLeader

Claude Fable 5.1 vs Grok 4

Verdict
  • Claude Fable 5.1 leads on quality: 70.7 vs 58.6.
  • Claude Fable 5.1 is stronger in agents & tools, coding, composite, knowledge, long context.
  • Grok 4 is stronger in human preference, instruction following, maths, multimodal, reasoning.
  • Grok 4 is 13× cheaper ($1.56 vs $20.00 per 1M blended).
MetricClaude Fable 5.1Grok 4
BenchLeader Index70.758.6
Agents & tools score72.155.3
Coding score68.561.2
Composite score95.057.4
Knowledge score70.558.3
Long context score69.369.0
Human preference score59.8
Instruction following score62.4
Maths score58.0
Multimodal score53.8
Reasoning score58.3
Blended price $/M$20.00$1.56
Output speed65 tok/s
Time to first answer295.0 s
Context window1M256k
GPQA Diamond87.0%
OTIS Mock AIME84.0%
Terminal-Bench27.2%
SimpleBench60.5%
Fiction.LiveBench 120k96.9%
Cybench43.0%
WeirdML45.7%
APEX-Agents47.4%15.2%
Epoch Capabilities Index164.2146.4
LMArena Text1411
LMArena Hard Prompts1420
LMArena Coding1435
LMArena Vision1210
AA Intelligence Index53.422.5
IFBench53.7%
AA-LCR85.3%68.0%
MMMU-Pro68.8%
AA-Omniscience43.52.1
Terminal-Bench Hard37.9%
GPQA Diamond (AA)93.7%87.7%
Humanity's Last Exam (AA)59.1%26.7%
SciCode (AA)63.1%
τ²-Bench Telecom (AA)74.8%
AIME (Vals)90.6%
LiveCodeBench90.5%83.3%
MMLU-Pro92.4%85.3%
IOI90.8%
LegalBench88.5%83.2%
CorpFin66.0%
TaxEval76.0%65.1%
MedQA92.5%
MGSM90.9%
Terminal-Bench 2.1 (Vals)85.0%
SWE-bench (Vals)57.8%
GPQA Diamond (Vals)93.4%88.1%
Vals Index68.8
IMO 202521.4%
MathArena Apex2.1%
MCP Atlas87.2%
PRBench Finance50.8%
PRBench Legal51.6%
HiL-Bench61.5%
Kagi LLM Benchmark73.6%
IFEval (HELM)94.9%
Omni-MATH (HELM)60.3%
WildBench (HELM)79.7%
MMLU-Pro (HELM)85.1%
GPQA Diamond (HELM)72.6%
HELM Capabilities mean78.5%
Aider Polyglot79.6%
ARC-AGI-179.6%
ARC-AGI-229.4%
BFCL Overall63.0%

Data as of 2026-09-12. Best configuration of each model; every score links to its source on the model pages.

Claude Fable 5.1 vs Grok 4: questions

Is Claude Fable 5.1 better than Grok 4?
Claude Fable 5.1 leads on quality: 70.7 vs 58.6. The BenchLeader Index combines every independent quality benchmark; Claude Fable 5.1 is ahead overall as of 2026-09-12, but check the category scores for your use.
Is Claude Fable 5.1 better than Grok 4 for coding?
Claude Fable 5.1 scores higher in coding (69 vs 61 on the category index, where 50 is average).
Is Claude Fable 5.1 better than Grok 4 for agentic tasks?
Claude Fable 5.1 scores higher in agentic tasks (72 vs 55 on the category index, where 50 is average).
Which is cheaper, Claude Fable 5.1 or Grok 4?
Grok 4 is cheaper: $1.56 against $20.00 per million tokens, blended at three input tokens per output token.
Which has the larger context window?
Claude Fable 5.1 accepts more context: 1M against 256k tokens.