MedCode
Assigning medical billing codes from clinical notes. Run by Vals AI.
As of 19 Sept 2026, Claude Opus 5 leads MedCode on BenchLeader with 63.6%, ahead of Gemini 3.1 Pro at 59.1%, across 90 model configurations with a published result.
- Published by
- Vals AI
- Category
- Knowledge
- Index weight
- Reference only
- Models
- 90
- Data as of
- 19 Sept 2026
Vals AI (vals.ai).
What the test looks like
Clinical notes are given and the model must assign the correct medical billing codes, a task with strict, checkable answers.
How it is scored
Accuracy, run by Vals AI.
What to keep in mind
Narrow but economically real; US coding systems only.
- 1Claude Opus 563.6%
- 2Gemini 3.1 Pro (high)59.1%
- 3Claude Fable 556.1%
- 4Gemini 3 Flash (high)55.9%
- 5Gemini 3.5 Flash (high)55.8%
- 6Claude Opus 4.754.9%
- 7Claude Fable 5.153.5%
- 8Gemini 3.7 Flash (high)53.4%
- 9Claude Opus 4.853.2%
- 10Gemini 3.6 Flash (high)53.1%
- 11GPT-5.1 (high)52.7%
- 12Gemini 3 Pro (high)52.2%
- 13Muse Spark51.3%
- 14Gemini 2.5 Pro50.6%
- 15GPT-5.2 (xhigh)49.8%
90 of 90
| # | ||||
|---|---|---|---|---|
| 1 | 63.6% | 66.7 | 2026-09-15 | |
| 2 | 59.1% | 60.7 | 2026-09-15 | |
| 3 | 56.1% | 68.3 | 2026-09-15 | |
| 4 | 55.9% | 58.7 | 2026-09-15 | |
| 5 | 55.8% | 63.6 | 2026-09-15 | |
| 6 | 54.9% | 64.5 | 2026-09-15 | |
| 7 | 53.5% | 64.5 | 2026-09-15 | |
| 8 | 53.4% | 64.6 | 2026-09-15 | |
| 9 | 53.2% | 61.9 | 2026-09-15 | |
| 10 | 53.1% | 61.7 | 2026-09-15 | |
| 11 | 52.7% | 58.7 | 2026-09-15 | |
| 12 | 52.2% | 61.1 | 2026-09-15 | |
| 13 | 51.3% | 65.7 | 2026-09-15 | |
| 14 | 50.6% | 54.5 | 2026-09-15 | |
| 15 | 49.8% | 62.0 | 2026-09-15 | |
| 16 | 49.6% | 58.6 | 2026-09-15 | |
| 17 | 49.4% | 64.1 | 2026-09-15 | |
| 18 | 49.2% | 60.8 | 2026-09-15 | |
| 19 | 49.1% | 62.5 | 2026-09-15 | |
| 20 | 49.1% | 67.6 | 2026-09-15 | |
| 21 | 48.9% | 64.2 | 2026-09-15 | |
| 22 | 48.5% | 71.8 | 2026-09-15 | |
| 23 | 48.2% | 63.7 | 2026-09-15 | |
| 24 | 48.1% | 64.5 | 2026-09-15 | |
| 25 | 47.6% | 49.1 | 2026-09-15 | |
| 26 | 47.5% | 57.2 | 2026-09-15 | |
| 27 | 47.3% | 53.0 | 2026-09-15 | |
| 28 | 47.2% | 57.2 | 2026-09-15 | |
| 29 | 46.3% | 57.5 | 2026-09-15 | |
| 30 | 45.2% | 58.2 | 2026-09-15 | |
| 31 | 44.7% | 64.5 | 2026-09-15 | |
| 32 | 44.1% | 53.4 | 2026-09-15 | |
| 33 | 44.0% | 68.8 | 2026-09-15 | |
| 34 | 43.5% | 44.3 | 2026-09-15 | |
| 35 | 43.4% | 64.3 | 2026-09-15 | |
| 36 | 43.3% | 61.0 | 2026-09-15 | |
| 37 | 43.0% | 55.3 | 2026-09-15 | |
| 38 | 42.9% | 65.8 | 2026-09-15 | |
| 39 | 42.5% | 56.0 | 2026-09-15 | |
| 40 | 42.4% | 60.2 | 2026-09-15 | |
| 41 | 41.6% | 56.9 | 2026-09-15 | |
| 42 | 41.4% | 55.0 | 2026-09-15 | |
| 43 | 41.4% | 51.9 | 2026-09-15 | |
| 44 | 41.3% | 65.4 | 2026-09-15 | |
| 45 | 41.2% | – | 2026-09-15 | |
| 46 | 41.0% | 49.1 | 2026-09-15 | |
| 47 | 40.8% | 52.1 | 2026-09-15 | |
| 48 | 40.7% | 66.5 | 2026-09-15 | |
| 49 | 40.6% | 54.3 | 2026-09-15 | |
| 50 | 40.5% | 52.1 | 2026-09-15 | |
| 51 | 40.5% | 64.2 | 2026-09-15 | |
| 52 | 40.4% | 48.4 | 2026-09-15 | |
| 53 | 40.3% | 51.6 | 2026-09-15 | |
| 54 | 40.1% | 60.5 | 2026-09-15 | |
| 55 | 39.3% | 58.5 | 2026-09-15 | |
| 56 | 38.8% | 61.0 | 2026-09-15 | |
| 57 | 38.6% | 48.1 | 2026-09-15 | |
| 58 | 38.4% | 52.5 | 2026-09-15 | |
| 59 | 38.1% | 58.5 | 2026-09-15 | |
| 60 | 38.1% | 50.5 | 2026-09-15 |
Cite as: BenchLeader, “MedCode leaderboard”, https://www.benchleader.com/benchmarks/vals_medcode, data as of 19 Sept 2026.
MedCode: questions
- What does MedCode measure?
- Clinical notes are given and the model must assign the correct medical billing codes, a task with strict, checkable answers. Scores are reported in percent of tasks solved; higher is better.
- Which AI model leads MedCode?
- Claude Opus 5 leads MedCode with 63.6% as of 19 Sept 2026, ahead of Gemini 3.1 Pro at 59.1%.
- How many models have MedCode results?
- 90 model configurations have a MedCode result on BenchLeader, all taken from Vals AI.
- Who runs MedCode and how often is it updated?
- MedCode is published by Vals AI. BenchLeader re-reads the published results every morning and records the date each result was published.
- Does MedCode count toward the BenchLeader Index?
- No. MedCode is shown for reference but left out of the composite index.