MedScribe
Doctor administrative paperwork from encounter notes. Run by Vals AI.
As of 19 Sept 2026, Claude Fable 5.1 leads MedScribe on BenchLeader with 91.3%, ahead of Claude Opus 5 at 91.0%, across 92 model configurations with a published result.
- Published by
- Vals AI
- Category
- Knowledge
- Index weight
- Reference only
- Models
- 92
- Data as of
- 19 Sept 2026
Vals AI (vals.ai).
What the test looks like
From an encounter transcript the model drafts the paperwork a doctor's office produces, graded against reference documents.
How it is scored
Rubric score, run by Vals AI.
What to keep in mind
Rubric-graded writing; style choices affect scores.
- 1Claude Fable 5.191.3%
- 2Claude Opus 591.0%
- 3Muse Spark 1.2 (xhigh)90.1%
- 4GLM-5.3-Flash (max)88.9%
- 5Muse Spark 1.1 (xhigh)88.9%
- 6GLM-5.3 (max)88.8%
- 7Claude Fable 588.5%
- 8GPT-5.1 (high)88.1%
- 9Kimi K388.0%
- 10GPT-6 Astra (max)87.9%
- 11MiniMax-M387.3%
- 12Grok 4.5 (high)86.9%
- 13GPT-5.5 (xhigh)86.9%
- 14Claude Opus 4.686.7%
- 15Grok 4.6 (high)86.5%
92 of 92
| # | ||||
|---|---|---|---|---|
| 1 | 91.3% | 64.5 | 2026-09-15 | |
| 2 | 91.0% | 66.7 | 2026-09-15 | |
| 3 | 90.1% | 64.1 | 2026-09-15 | |
| 4 | 88.9% | 55.3 | 2026-09-15 | |
| 5 | 88.9% | 61.8 | 2026-09-15 | |
| 6 | 88.8% | 65.8 | 2026-09-15 | |
| 7 | 88.5% | 68.3 | 2026-09-15 | |
| 8 | 88.1% | 58.7 | 2026-09-15 | |
| 9 | 88.0% | 64.2 | 2026-09-15 | |
| 10 | 87.9% | 71.8 | 2026-09-15 | |
| 11 | 87.3% | 57.5 | 2026-09-15 | |
| 12 | 86.9% | 61.0 | 2026-09-15 | |
| 13 | 86.9% | 67.6 | 2026-09-15 | |
| 14 | 86.7% | 63.7 | 2026-09-15 | |
| 15 | 86.5% | 64.5 | 2026-09-15 | |
| 16 | 86.1% | 62.5 | 2026-09-15 | |
| 17 | 85.9% | 65.7 | 2026-09-15 | |
| 18 | 85.8% | 61.9 | 2026-09-15 | |
| 19 | 85.5% | – | 2026-09-15 | |
| 20 | 85.3% | 60.8 | 2026-09-15 | |
| 21 | 85.2% | 68.8 | 2026-09-15 | |
| 22 | 85.2% | 47.5 | 2026-09-15 | |
| 23 | 85.0% | 66.5 | 2026-09-15 | |
| 24 | 84.5% | 54.3 | 2026-09-15 | |
| 25 | 84.5% | 64.5 | 2026-09-15 | |
| 26 | 84.4% | 62.0 | 2026-09-15 | |
| 27 | 84.4% | 60.2 | 2026-09-15 | |
| 28 | Inkling SmallThinking Machinesopen ↗ | 84.1% | 56.3 | 2026-09-15 |
| 29 | 84.1% | 53.4 | 2026-09-15 | |
| 30 | 83.9% | 64.6 | 2026-09-15 | |
| 31 | 83.8% | 59.2 | 2026-09-15 | |
| 32 | 83.7% | 59.2 | 2026-09-15 | |
| 33 | 83.7% | 58.6 | 2026-09-15 | |
| 34 | 83.5% | 52.1 | 2026-09-15 | |
| 35 | 83.3% | 58.2 | 2026-09-15 | |
| 36 | 83.0% | 48.4 | 2026-09-15 | |
| 37 | 83.0% | 64.5 | 2026-09-15 | |
| 38 | 82.9% | 64.3 | 2026-09-15 | |
| 39 | 82.9% | 52.5 | 2026-09-15 | |
| 40 | 81.6% | 52.6 | 2026-09-15 | |
| 41 | 80.9% | 53.6 | 2026-09-15 | |
| 42 | 80.8% | 55.1 | 2026-09-15 | |
| 43 | 80.6% | 55.3 | 2026-09-15 | |
| 44 | 80.4% | 55.0 | 2026-09-15 | |
| 45 | 80.2% | 56.0 | 2026-09-15 | |
| 46 | 79.9% | 56.1 | 2026-09-15 | |
| 47 | 79.7% | 41.1 | 2026-09-15 | |
| 48 | 79.7% | 61.7 | 2026-09-15 | |
| 49 | 79.4% | 61.0 | 2026-09-15 | |
| 50 | 78.7% | 56.0 | 2026-09-15 | |
| 51 | 78.5% | 51.6 | 2026-09-15 | |
| 52 | 78.2% | 60.5 | 2026-09-15 | |
| 53 | 78.2% | 58.5 | 2026-09-15 | |
| 54 | 78.0% | 52.1 | 2026-09-15 | |
| 55 | 77.5% | 65.4 | 2026-09-15 | |
| 56 | 77.5% | 43.0 | 2026-09-15 | |
| 57 | 77.1% | – | 2026-09-15 | |
| 58 | 77.1% | 49.1 | 2026-09-15 | |
| 59 | 77.0% | 58.5 | 2026-09-15 | |
| 60 | 76.7% | 53.0 | 2026-09-15 |
Cite as: BenchLeader, “MedScribe leaderboard”, https://www.benchleader.com/benchmarks/vals_medscribe, data as of 19 Sept 2026.
MedScribe: questions
- What does MedScribe measure?
- From an encounter transcript the model drafts the paperwork a doctor's office produces, graded against reference documents. Scores are reported in percent of tasks solved; higher is better.
- Which AI model leads MedScribe?
- Claude Fable 5.1 leads MedScribe with 91.3% as of 19 Sept 2026, ahead of Claude Opus 5 at 91.0%.
- How many models have MedScribe results?
- 92 model configurations have a MedScribe result on BenchLeader, all taken from Vals AI.
- Who runs MedScribe and how often is it updated?
- MedScribe is published by Vals AI. BenchLeader re-reads the published results every morning and records the date each result was published.
- Does MedScribe count toward the BenchLeader Index?
- No. MedScribe is shown for reference but left out of the composite index.