BenchLeader

Claude Opus 4.1 vs GPT-5.4 Pro

Verdict
  • GPT-5.4 Pro leads on quality: 64.3 vs 57.2.
  • Claude Opus 4.1 (thinking) is stronger in coding, human preference, maths.
  • GPT-5.4 Pro is stronger in instruction following, knowledge, multimodal, reasoning.
  • Claude Opus 4.1 (thinking) is 2.3× cheaper ($30.00 vs $67.50 per 1M blended).
  • Claude Opus 4.1 (thinking) streams 10.5× faster (11 vs 1 tokens per second).
MetricClaude Opus 4.1 (thinking)GPT-5.4 Pro
BenchLeader Index57.264.3
Coding score53.3
Human preference score64.5
Instruction following score50.467.3
Knowledge score58.373.6
Maths score59.3
Multimodal score60.969.8
Reasoning score65.675.3
Blended price $/M$30.00$67.50
Output speed11 tok/s1 tok/s
Time to first answer3.2 s6.3 s
Context window200k1.1M
Humanity's Last Exam44.3%
SimpleBench74.1%
Epoch Capabilities Index158.9
LMArena Text1450
LMArena Hard Prompts1480
LMArena Coding1512
AIME (Vals)78.2%
LiveCodeBench66.5%
MMLU-Pro87.9%
TaxEval73.7%
MedQA93.6%
MGSM94.4%
GPQA Diamond (Vals)76.3%
MultiChallenge57.2%69.2%
VISTA48.4%53.9%
MultiNRC38.4%62.3%
TutorBench50.8%56.6%

Data as of 2026-09-13. Best configuration of each model; every score links to its source on the model pages.

Claude Opus 4.1 vs GPT-5.4 Pro: questions

Is Claude Opus 4.1 better than GPT-5.4 Pro?
GPT-5.4 Pro leads on quality: 64.3 vs 57.2. The BenchLeader Index combines every independent quality benchmark; GPT-5.4 Pro is ahead overall as of 2026-09-13, but check the category scores for your use.
Which is cheaper, Claude Opus 4.1 or GPT-5.4 Pro?
Claude Opus 4.1 is cheaper: $30.00 against $67.50 per million tokens, blended at three input tokens per output token.
Which is faster, Claude Opus 4.1 or GPT-5.4 Pro?
Claude Opus 4.1 streams faster: 11 against 1 output tokens per second.
Which has the larger context window?
GPT-5.4 Pro accepts more context: 1.1M against 200k tokens.