BenchLeader

MedScribe

Doctor administrative paperwork from encounter notes. Run by Vals AI.

As of 19 Sept 2026, Claude Fable 5.1 leads MedScribe on BenchLeader with 91.3%, ahead of Claude Opus 5 at 91.0%, across 92 model configurations with a published result.

Published by
Vals AI
Category
Knowledge
Index weight
Reference only
Models
92
Data as of
19 Sept 2026

Vals AI (vals.ai).

What the test looks like

From an encounter transcript the model drafts the paperwork a doctor's office produces, graded against reference documents.

How it is scored

Rubric score, run by Vals AI.

What to keep in mind

Rubric-graded writing; style choices affect scores.

92 of 92
#
1Claude Fable 5.1Anthropic91.3%64.52026-09-15
2Claude Opus 5Anthropic91.0%66.72026-09-15
3Muse Spark 1.2xhighMeta90.1%64.12026-09-15
4GLM-5.3-FlashmaxZhipu AIopen ↗88.9%55.32026-09-15
5Muse Spark 1.1xhighMeta88.9%61.82026-09-15
6GLM-5.3maxZhipu AIopen ↗88.8%65.82026-09-15
7Claude Fable 5Anthropic88.5%68.32026-09-15
8GPT-5.1highOpenAI88.1%58.72026-09-15
9Kimi K3Moonshot AIopen ↗88.0%64.22026-09-15
10GPT-6 AstramaxOpenAI87.9%71.82026-09-15
11MiniMax-M3MiniMaxopen ↗87.3%57.52026-09-15
12Grok 4.5highxAI86.9%61.02026-09-15
13GPT-5.5xhighOpenAI86.9%67.62026-09-15
14Claude Opus 4.6Anthropic86.7%63.72026-09-15
15Grok 4.6highxAI86.5%64.52026-09-15
16Claude Opus 4.6thinkingAnthropic86.1%62.52026-09-15
17Muse SparkMeta85.9%65.72026-09-15
18Claude Opus 4.8Anthropic85.8%61.92026-09-15
19DeepSeek V4.1 FlashhighDeepSeekopen ↗85.5%2026-09-15
20Claude Opus 4.5thinkingAnthropic85.3%60.82026-09-15
21GPT-5.6 SolmaxOpenAI85.2%68.82026-09-15
22Claude Haiku 4.5thinkingAnthropic85.2%47.52026-09-15
23Qwen3 8maxAlibaba85.0%66.52026-09-15
24Claude Sonnet 4.5Anthropic84.5%54.32026-09-15
25Gemini 3.8 FlashhighGoogle84.5%64.52026-09-15
26GPT-5.2xhighOpenAI84.4%62.02026-09-15
27GPT-5.6 LunamaxOpenAI84.4%60.22026-09-15
28Inkling SmallThinking Machinesopen ↗84.1%56.32026-09-15
29Claude Sonnet 4.5thinkingAnthropic84.1%53.42026-09-15
30Gemini 3.7 FlashhighGoogle83.9%64.62026-09-15
31Qwen3.8 27BxhighAlibabaopen ↗83.8%59.22026-09-15
32MiMo-V2.5-ProXiaomiopen ↗83.7%59.22026-09-15
33GPT-5highOpenAI83.7%58.62026-09-15
34GLM-5.2Zhipu AIopen ↗83.5%52.12026-09-15
35Claude Opus 4.5Anthropic83.3%58.22026-09-15
36Gemini 2.5 FlashthinkingGoogle83.0%48.42026-09-15
37Claude Opus 4.7Anthropic83.0%64.52026-09-15
38GPT-5.6 TerraxhighOpenAI82.9%64.32026-09-15
39Gemini 2.5 FlashGoogle82.9%52.52026-09-15
40Grok 4 FastthinkingxAI81.6%52.62026-09-15
41Ling 3.0 FlashAnt Groupopen ↗80.9%53.62026-09-15
42MiniMax-M2MiniMaxopen80.8%55.12026-09-15
43GPT-5 minihighOpenAI80.6%55.32026-09-15
44DeepSeek V4 Flash 0731highDeepSeekopen ↗80.4%55.02026-09-15
45DeepSeek V4 Pro 0813maxDeepSeekopen ↗80.2%56.02026-09-15
46MiniMax-M2.7MiniMaxopen ↗79.9%56.12026-09-15
47Grok 4 Fastno reasoningxAI79.7%41.12026-09-15
48Gemini 3.6 FlashhighGoogle79.7%61.72026-09-15
49Qwen3 7maxAlibaba79.4%61.02026-09-15
50Grok 4.1thinkingxAI78.7%56.02026-09-15
51Gemini 2.5 Flash 09 2025thinkingGoogle78.5%51.62026-09-15
52Kimi K2.6Moonshot AIopen ↗78.2%60.52026-09-15
53Grok 4xAI78.2%58.52026-09-15
54Gemini 2.5 Flash 09 2025Google78.0%52.12026-09-15
55GPT-5.4xhighOpenAI77.5%65.42026-09-15
56Grok 4.1no reasoningxAI77.5%43.02026-09-15
57Qwen3-VL PlusAlibaba77.1%2026-09-15
58GPT-5.4 nanohighOpenAI77.1%49.12026-09-15
59Qwen3.6 PlusAlibabaopen77.0%58.52026-09-15
60o3highOpenAI76.7%53.02026-09-15

Cite as: BenchLeader, “MedScribe leaderboard”, https://www.benchleader.com/benchmarks/vals_medscribe, data as of 19 Sept 2026.

MedScribe: questions

What does MedScribe measure?
From an encounter transcript the model drafts the paperwork a doctor's office produces, graded against reference documents. Scores are reported in percent of tasks solved; higher is better.
Which AI model leads MedScribe?
Claude Fable 5.1 leads MedScribe with 91.3% as of 19 Sept 2026, ahead of Claude Opus 5 at 91.0%.
How many models have MedScribe results?
92 model configurations have a MedScribe result on BenchLeader, all taken from Vals AI.
Who runs MedScribe and how often is it updated?
MedScribe is published by Vals AI. BenchLeader re-reads the published results every morning and records the date each result was published.
Does MedScribe count toward the BenchLeader Index?
No. MedScribe is shown for reference but left out of the composite index.