METR Time Horizons
METR's task suite behind the 50%-time-horizon measure.
As of 19 Sept 2026, Claude Mythos Preview leads METR Time Horizons on BenchLeader with 85.2%, ahead of Claude Opus 4.6 at 78.9%, across 42 model configurations with a published result.
- Published by
- METRdata via Epoch AI Benchmarking Hub
- Category
- Agents & tools
- Index weight
- Reference only
- Models
- 42
- Data as of
- 19 Sept 2026
CC BY 4.0 — Epoch AI, ‘AI Benchmarking Hub’, epoch.ai/benchmarks. Mirrored boards credit their original publishers.
What the test looks like
Software and research tasks of increasing length; METR fits the task duration at which a model succeeds half the time.
How it is scored
Average success across the task suite, as published by METR.
What to keep in mind
The headline METR number is a fitted time horizon; this is the underlying success rate.
- 1Claude Mythos Preview85.2%
- 2Claude Opus 4.678.9%
- 3Gemini 3.1 Pro77.0%
- 4GPT-5.2 (high)75.3%
- 5Claude Opus 4.575.0%
- 6GPT-5.3 Codex74.5%
- 7GPT-5.4 (xhigh)74.3%
- 8GPT-5.474.3%
- 9Gemini 3 Pro71.0%
- 10GPT-5.1-Codex-Max (max)70.8%
- 11GPT-5 (medium)69.6%
- 12GPT-5 (high)69.4%
- 13Claude Sonnet 4.567.4%
- 14Claude Opus 4.166.8%
- 15Grok 466.6%
| # | ||||
|---|---|---|---|---|
| 1 | 85.2% | – | 2026-04-07 | |
| 2 | 78.9% | 63.7 | 2026-02-05 | |
| 3 | 77.0% | 63.9 | 2026-02-19 | |
| 4 | 75.3% | 55.7 | 2025-12-11 | |
| 5 | 75.0% | 58.2 | 2025-11-24 | |
| 6 | 74.5% | 54.2 | 2026-02-05 | |
| 7 | 74.3% | 65.4 | 2026-03-05 | |
| 8 | 74.3% | 59.2 | 2026-03-05 | |
| 9 | 71.0% | 61.1 | 2025-11-18 | |
| 10 | 70.8% | – | 2025-11-19 | |
| 11 | 69.6% | 58.5 | 2025-08-07 | |
| 12 | 69.4% | 58.6 | 2025-08-07 | |
| 13 | 67.4% | 54.3 | 2025-09-29 | |
| 14 | 66.8% | 51.9 | 2025-08-05 | |
| 15 | 66.6% | 58.5 | 2025-07-09 | |
| 16 | 65.4% | 53.7 | 2025-04-16 | |
| 17 | 63.9% | 52.9 | 2025-05-22 | |
| 18 | 63.9% | 52.4 | 2025-04-16 | |
| 19 | 63.6% | 60.3 | 2025-04-16 | |
| 20 | 62.0% | 49.7 | 2025-05-22 | |
| 21 | 60.0% | 50.2 | 2025-02-24 | |
| 22 | 59.2% | 52.2 | 2025-11-06 | |
| 23 | 56.6% | 46.9 | 2025-08-05 | |
| 24 | 55.4% | 54.5 | 2025-06-05 | |
| 25 | 53.8% | 50.2 | 2025-05-28 | |
| 26 | 51.9% | 48.7 | 2025-01-20 | |
| 27 | 51.0% | 48.4 | 2024-12-17 | |
| 28 | 47.4% | 44.8 | 2024-12-26 | |
| 29 | 45.2% | 46.6 | 2024-10-22 | |
| 30 | 45.1% | 53.8 | 2024-09-12 | |
| 31 | 40.8% | 44.3 | 2024-11-20 | |
| 32 | 36.7% | 41.5 | 2024-04-09 | |
| 33 | 36.1% | 43.9 | 2023-03-14 | |
| 34 | 35.8% | 42.7 | 2024-09-19 | |
| 35 | 35.2% | 46.0 | 2024-01-25 | |
| 36 | 29.9% | 41.3 | 2024-06-07 | |
| 37 | 29.5% | 39.6 | 2024-02-29 | |
| 38 | 29.3% | 40.2 | 2023-06-13 | |
| 39 | 28.9% | 44.7 | 2023-11-06 | |
| 40 | 21.5% | – | 2023-09-18 | |
| 41 | 16.2% | – | – | |
| 42 | 10.1% | – | 2019-11-05 |
Cite as: BenchLeader, “METR Time Horizons leaderboard”, https://www.benchleader.com/benchmarks/metr_time_horizons, data as of 19 Sept 2026.
METR Time Horizons: questions
- What does METR Time Horizons measure?
- Software and research tasks of increasing length; METR fits the task duration at which a model succeeds half the time. Scores are reported in percent of tasks solved; higher is better.
- Which AI model leads METR Time Horizons?
- Claude Mythos Preview leads METR Time Horizons with 85.2% as of 19 Sept 2026, ahead of Claude Opus 4.6 at 78.9%.
- How many models have METR Time Horizons results?
- 42 model configurations have a METR Time Horizons result on BenchLeader, all taken from METR via Epoch AI Benchmarking Hub.
- Who runs METR Time Horizons and how often is it updated?
- METR Time Horizons is published by METR. BenchLeader re-reads the published results every morning and records the date each result was published.
- Does METR Time Horizons count toward the BenchLeader Index?
- No. METR Time Horizons is shown for reference but left out of the composite index.