BenchLeader

METR Time Horizons

METR's task suite behind the 50%-time-horizon measure.

As of 19 Sept 2026, Claude Mythos Preview leads METR Time Horizons on BenchLeader with 85.2%, ahead of Claude Opus 4.6 at 78.9%, across 42 model configurations with a published result.

Published by
METRdata via Epoch AI Benchmarking Hub
Category
Agents & tools
Index weight
Reference only
Models
42
Data as of
19 Sept 2026

CC BY 4.0 — Epoch AI, ‘AI Benchmarking Hub’, epoch.ai/benchmarks. Mirrored boards credit their original publishers.

What the test looks like

Software and research tasks of increasing length; METR fits the task duration at which a model succeeds half the time.

How it is scored

Average success across the task suite, as published by METR.

What to keep in mind

The headline METR number is a fitted time horizon; this is the underlying success rate.

42 of 42
#
1Claude Mythos PreviewAnthropic85.2%2026-04-07
2Claude Opus 4.6Anthropic78.9%63.72026-02-05
3Gemini 3.1 ProGoogle77.0%63.92026-02-19
4GPT-5.2highOpenAI75.3%55.72025-12-11
5Claude Opus 4.5Anthropic75.0%58.22025-11-24
6GPT-5.3 CodexOpenAI74.5%54.22026-02-05
7GPT-5.4xhighOpenAI74.3%65.42026-03-05
8GPT-5.4OpenAI74.3%59.22026-03-05
9Gemini 3 ProGoogle71.0%61.12025-11-18
10GPT-5.1-Codex-MaxmaxOpenAI70.8%2025-11-19
11GPT-5mediumOpenAI69.6%58.52025-08-07
12GPT-5highOpenAI69.4%58.62025-08-07
13Claude Sonnet 4.5Anthropic67.4%54.32025-09-29
14Claude Opus 4.1Anthropic66.8%51.92025-08-05
15Grok 4xAI66.6%58.52025-07-09
16o3mediumOpenAI65.4%53.72025-04-16
17Claude Opus 4Anthropic63.9%52.92025-05-22
18o4-minimediumOpenAI63.9%52.42025-04-16
19o3OpenAI63.6%60.32025-04-16
20Claude Sonnet 4Anthropic62.0%49.72025-05-22
21Claude 3.7 SonnetAnthropic60.0%50.22025-02-24
22Kimi K2 ThinkingthinkingMoonshot AIopen59.2%52.22025-11-06
23gpt-oss-120bOpenAIopen56.6%46.92025-08-05
24Gemini 2.5 ProGoogle55.4%54.52025-06-05
25DeepSeek R1 0528DeepSeekopen53.8%50.22025-05-28
26DeepSeek R1DeepSeekopen51.9%48.72025-01-20
27o1mediumOpenAI51.0%48.42024-12-17
28DeepSeek V3DeepSeekopen47.4%44.82024-12-26
29Claude 3.5 SonnetAnthropic45.2%46.62024-10-22
30o1OpenAI45.1%53.82024-09-12
31GPT-4oOpenAI40.8%44.32024-11-20
32GPT-4 TurboOpenAI36.7%41.52024-04-09
33GPT 4 0314OpenAI36.1%43.92023-03-14
34Qwen2.5 72BAlibabaopen35.8%42.72024-09-19
35GPT 4 0125OpenAI35.2%46.02024-01-25
36Qwen2-72BAlibabaopen ↗29.9%41.32024-06-07
37Claude 3 OpusAnthropic29.5%39.62024-02-29
38GPT 4 0613OpenAI29.3%40.22023-06-13
39GPT-4OpenAI28.9%44.72023-11-06
40GPT-3.5-turboOpenAI21.5%2023-09-18
41Davinci 002OpenAI16.2%
42Gpt2 XlOpenAI10.1%2019-11-05

Cite as: BenchLeader, “METR Time Horizons leaderboard”, https://www.benchleader.com/benchmarks/metr_time_horizons, data as of 19 Sept 2026.

METR Time Horizons: questions

What does METR Time Horizons measure?
Software and research tasks of increasing length; METR fits the task duration at which a model succeeds half the time. Scores are reported in percent of tasks solved; higher is better.
Which AI model leads METR Time Horizons?
Claude Mythos Preview leads METR Time Horizons with 85.2% as of 19 Sept 2026, ahead of Claude Opus 4.6 at 78.9%.
How many models have METR Time Horizons results?
42 model configurations have a METR Time Horizons result on BenchLeader, all taken from METR via Epoch AI Benchmarking Hub.
Who runs METR Time Horizons and how often is it updated?
METR Time Horizons is published by METR. BenchLeader re-reads the published results every morning and records the date each result was published.
Does METR Time Horizons count toward the BenchLeader Index?
No. METR Time Horizons is shown for reference but left out of the composite index.