BenchLeader

MLCR

Medical long-context reasoning: clinical questions answered over long patient records, scored on accuracy, completeness and concision together.

As of 22 Sept 2026, Claude Fable 5.1 leads MLCR on BenchLeader with 71.1%, ahead of Claude Fable 5 at 64.4%, across 34 model configurations with a published result.

Published by
Artificial Analysis
Category
Long context
Index weight
Reference only
Models
34
Data as of
22 Sept 2026

Source: Artificial Analysis (artificialanalysis.ai). Data taken from the public leaderboard.

34 of 34
#
1Claude Fable 5.1thinkingAnthropic71.1%71.02026-09-01
2Claude Fable 5thinkingAnthropic64.4%70.32026-06-09
3Claude Opus 5highAnthropic59.4%69.82026-07-24
4Claude Opus 5xhighAnthropic58.3%69.32026-07-24
5Claude Opus 5mediumAnthropic56.1%66.42026-07-24
6Claude Opus 5maxAnthropic55.6%69.52026-07-24
7Claude Sonnet 5maxAnthropic55.0%60.02026-06-30
8GLM 5.3 FlashZhipu AIopen ↗51.1%63.62026-08-26
9GLM 5.3maxZhipu AIopen ↗48.3%65.42026-08-18
10Muse Spark 1.3maxMeta43.3%69.22026-09-02
11Kimi K3maxMoonshot AIopen ↗38.3%66.72026-07-16
12GPT-6 AstramaxOpenAI35.0%71.72026-09-03
13GPT-5.6 TerramaxOpenAI31.7%64.82026-07-09
14Muse Spark 1.2xhighMeta31.1%63.72026-08-05
15GPT-5.6 SolmaxOpenAI26.1%68.52026-07-09
16DeepSeek V4.1 FlashmaxDeepSeekopen ↗22.8%61.42026-09-10
17Gemini 3.8 FlashhighGoogle21.7%64.32026-09-02
18Qwen3.8 27BxhighAlibabaopen ↗21.7%58.92026-08-14
19Qwen3.8 Max (0902)maxAlibaba20.0%66.22026-09-02
20Muse GlimmerhighMetaopen ↗20.0%52.52026-08-10
21GPT-5.6 LunamaxOpenAI19.4%59.92026-07-09
22Gemini 3.5 FlashhighGoogle18.3%63.42026-05-19
23DeepSeek V4 PromaxDeepSeekopen ↗17.8%63.72026-08-13
24MiniMax M3MiniMaxopen ↗17.2%57.22026-06-01
25NewStep 5 PreviewStepFun16.7%64.82026-09-18
26NewGrok 4.7xhighSpaceXAI15.0%61.92026-09-21
27Grok 4.6highSpaceXAI12.2%64.22026-08-12
28Nemotron 3 Ultra 550B A55BthinkingNVIDIAopen ↗11.1%57.32026-06-04
29Gemini 3.5 Flash LiteGoogle7.2%55.22026-07-21
30Nemotron 3 Super 120B A12bthinkingNVIDIAopen ↗3.3%51.22026-03-11
31Mistral Medium 3.5Mistral AIopen1.7%51.12026-04-29
32gpt-oss-120bhighOpenAIopen ↗1.1%49.22025-08-05
33Nemotron 3.5 LightningNVIDIAopen ↗1.1%48.02026-08-11
34Qwen3.8 2.4T A95BAlibabaopen ↗0.0%64.52026-08-12

Cite as: BenchLeader, “MLCR leaderboard”, https://www.benchleader.com/benchmarks/aa_mlcr, data as of 22 Sept 2026.

MLCR: questions

What does MLCR measure?
Medical long-context reasoning: clinical questions answered over long patient records, scored on accuracy, completeness and concision together. Scores are reported in percent of tasks solved; higher is better.
Which AI model leads MLCR?
Claude Fable 5.1 leads MLCR with 71.1% as of 22 Sept 2026, ahead of Claude Fable 5 at 64.4%.
How many models have MLCR results?
34 model configurations have a MLCR result on BenchLeader, all taken from Artificial Analysis.
Who runs MLCR and how often is it updated?
MLCR is published by Artificial Analysis. BenchLeader re-reads the published results every morning and records the date each result was published.
Does MLCR count toward the BenchLeader Index?
No. MLCR is shown for reference but left out of the composite index.