MLCR
Medical long-context reasoning: clinical questions answered over long patient records, scored on accuracy, completeness and concision together.
As of 22 Sept 2026, Claude Fable 5.1 leads MLCR on BenchLeader with 71.1%, ahead of Claude Fable 5 at 64.4%, across 34 model configurations with a published result.
- Published by
- Artificial Analysis
- Category
- Long context
- Index weight
- Reference only
- Models
- 34
- Data as of
- 22 Sept 2026
Source: Artificial Analysis (artificialanalysis.ai). Data taken from the public leaderboard.
- 1Claude Fable 5.1 (thinking)71.1%
- 2Claude Fable 5 (thinking)64.4%
- 3Claude Opus 5 (high)59.4%
- 4Claude Opus 5 (xhigh)58.3%
- 5Claude Opus 5 (medium)56.1%
- 6Claude Opus 5 (max)55.6%
- 7Claude Sonnet 5 (max)55.0%
- 8GLM 5.3 Flash51.1%
- 9GLM 5.3 (max)48.3%
- 10Muse Spark 1.3 (max)43.3%
- 11Kimi K3 (max)38.3%
- 12GPT-6 Astra (max)35.0%
- 13GPT-5.6 Terra (max)31.7%
- 14Muse Spark 1.2 (xhigh)31.1%
- 15GPT-5.6 Sol (max)26.1%
34 of 34
| # | ||||
|---|---|---|---|---|
| 1 | 71.1% | 71.0 | 2026-09-01 | |
| 2 | 64.4% | 70.3 | 2026-06-09 | |
| 3 | 59.4% | 69.8 | 2026-07-24 | |
| 4 | 58.3% | 69.3 | 2026-07-24 | |
| 5 | 56.1% | 66.4 | 2026-07-24 | |
| 6 | 55.6% | 69.5 | 2026-07-24 | |
| 7 | 55.0% | 60.0 | 2026-06-30 | |
| 8 | 51.1% | 63.6 | 2026-08-26 | |
| 9 | 48.3% | 65.4 | 2026-08-18 | |
| 10 | 43.3% | 69.2 | 2026-09-02 | |
| 11 | 38.3% | 66.7 | 2026-07-16 | |
| 12 | 35.0% | 71.7 | 2026-09-03 | |
| 13 | 31.7% | 64.8 | 2026-07-09 | |
| 14 | 31.1% | 63.7 | 2026-08-05 | |
| 15 | 26.1% | 68.5 | 2026-07-09 | |
| 16 | 22.8% | 61.4 | 2026-09-10 | |
| 17 | 21.7% | 64.3 | 2026-09-02 | |
| 18 | 21.7% | 58.9 | 2026-08-14 | |
| 19 | 20.0% | 66.2 | 2026-09-02 | |
| 20 | 20.0% | 52.5 | 2026-08-10 | |
| 21 | 19.4% | 59.9 | 2026-07-09 | |
| 22 | 18.3% | 63.4 | 2026-05-19 | |
| 23 | 17.8% | 63.7 | 2026-08-13 | |
| 24 | 17.2% | 57.2 | 2026-06-01 | |
| 25 | New | 16.7% | 64.8 | 2026-09-18 |
| 26 | New | 15.0% | 61.9 | 2026-09-21 |
| 27 | 12.2% | 64.2 | 2026-08-12 | |
| 28 | 11.1% | 57.3 | 2026-06-04 | |
| 29 | 7.2% | 55.2 | 2026-07-21 | |
| 30 | 3.3% | 51.2 | 2026-03-11 | |
| 31 | 1.7% | 51.1 | 2026-04-29 | |
| 32 | 1.1% | 49.2 | 2025-08-05 | |
| 33 | 1.1% | 48.0 | 2026-08-11 | |
| 34 | 0.0% | 64.5 | 2026-08-12 |
Cite as: BenchLeader, “MLCR leaderboard”, https://www.benchleader.com/benchmarks/aa_mlcr, data as of 22 Sept 2026.
MLCR: questions
- What does MLCR measure?
- Medical long-context reasoning: clinical questions answered over long patient records, scored on accuracy, completeness and concision together. Scores are reported in percent of tasks solved; higher is better.
- Which AI model leads MLCR?
- Claude Fable 5.1 leads MLCR with 71.1% as of 22 Sept 2026, ahead of Claude Fable 5 at 64.4%.
- How many models have MLCR results?
- 34 model configurations have a MLCR result on BenchLeader, all taken from Artificial Analysis.
- Who runs MLCR and how often is it updated?
- MLCR is published by Artificial Analysis. BenchLeader re-reads the published results every morning and records the date each result was published.
- Does MLCR count toward the BenchLeader Index?
- No. MLCR is shown for reference but left out of the composite index.