BenchLeader

MedCode

Assigning medical billing codes from clinical notes. Run by Vals AI.

As of 19 Sept 2026, Claude Opus 5 leads MedCode on BenchLeader with 63.6%, ahead of Gemini 3.1 Pro at 59.1%, across 90 model configurations with a published result.

Published by
Vals AI
Category
Knowledge
Index weight
Reference only
Models
90
Data as of
19 Sept 2026

Vals AI (vals.ai).

What the test looks like

Clinical notes are given and the model must assign the correct medical billing codes, a task with strict, checkable answers.

How it is scored

Accuracy, run by Vals AI.

What to keep in mind

Narrow but economically real; US coding systems only.

90 of 90
#
1Claude Opus 5Anthropic63.6%66.72026-09-15
2Gemini 3.1 ProhighGoogle59.1%60.72026-09-15
3Claude Fable 5Anthropic56.1%68.32026-09-15
4Gemini 3 FlashhighGoogle55.9%58.72026-09-15
5Gemini 3.5 FlashhighGoogle55.8%63.62026-09-15
6Claude Opus 4.7Anthropic54.9%64.52026-09-15
7Claude Fable 5.1Anthropic53.5%64.52026-09-15
8Gemini 3.7 FlashhighGoogle53.4%64.62026-09-15
9Claude Opus 4.8Anthropic53.2%61.92026-09-15
10Gemini 3.6 FlashhighGoogle53.1%61.72026-09-15
11GPT-5.1highOpenAI52.7%58.72026-09-15
12Gemini 3 ProhighGoogle52.2%61.12026-09-15
13Muse SparkMeta51.3%65.72026-09-15
14Gemini 2.5 ProGoogle50.6%54.52026-09-15
15GPT-5.2xhighOpenAI49.8%62.02026-09-15
16GPT-5highOpenAI49.6%58.62026-09-15
17Muse Spark 1.2xhighMeta49.4%64.12026-09-15
18Claude Opus 4.5thinkingAnthropic49.2%60.82026-09-15
19Claude Opus 4.6thinkingAnthropic49.1%62.52026-09-15
20GPT-5.5xhighOpenAI49.1%67.62026-09-15
21Kimi K3Moonshot AIopen ↗48.9%64.22026-09-15
22GPT-6 AstramaxOpenAI48.5%71.82026-09-15
23Claude Opus 4.6Anthropic48.2%63.72026-09-15
24Gemini 3.8 FlashhighGoogle48.1%64.52026-09-15
25Gemini 3.1 Flash LitehighGoogle47.6%49.12026-09-15
26Claude Sonnet 5Anthropic47.5%57.22026-09-15
27o3highOpenAI47.3%53.02026-09-15
28Claude Opus 4.1thinkingAnthropic47.2%57.22026-09-15
29MiniMax-M3MiniMaxopen ↗46.3%57.52026-09-15
30Claude Opus 4.5Anthropic45.2%58.22026-09-15
31Grok 4.6highxAI44.7%64.52026-09-15
32Claude Sonnet 4.5thinkingAnthropic44.1%53.42026-09-15
33GPT-5.6 SolmaxOpenAI44.0%68.82026-09-15
34Gemini 3.5 Flash LitehighGoogle43.5%44.32026-09-15
35GPT-5.6 TerraxhighOpenAI43.4%64.32026-09-15
36Grok 4.5highxAI43.3%61.02026-09-15
37GPT-5 minihighOpenAI43.0%55.32026-09-15
38GLM-5.3maxZhipu AIopen ↗42.9%65.82026-09-15
39DeepSeek V4 Pro 0813maxDeepSeekopen ↗42.5%56.02026-09-15
40GPT-5.6 LunamaxOpenAI42.4%60.22026-09-15
41GLM-5.1Zhipu AIopen ↗41.6%56.92026-09-15
42DeepSeek V4 Flash 0731highDeepSeekopen ↗41.4%55.02026-09-15
43Claude Opus 4.1Anthropic41.4%51.92026-09-15
44GPT-5.4xhighOpenAI41.3%65.42026-09-15
45DeepSeek V4.1 FlashhighDeepSeekopen ↗41.2%2026-09-15
46GPT-5.4 nanohighOpenAI41.0%49.12026-09-15
47GLM-5.2Zhipu AIopen ↗40.8%52.12026-09-15
48Qwen3 8maxAlibaba40.7%66.52026-09-15
49Claude Sonnet 4.5Anthropic40.6%54.32026-09-15
50Gemini 2.5 Flash 09 2025Google40.5%52.12026-09-15
51DeepSeek V4 PromaxDeepSeekopen ↗40.5%64.22026-09-15
52Gemini 2.5 FlashthinkingGoogle40.4%48.42026-09-15
53Gemini 2.5 Flash 09 2025thinkingGoogle40.3%51.62026-09-15
54Kimi K2.6Moonshot AIopen ↗40.1%60.52026-09-15
55Kimi K2.5thinkingMoonshot AIopen39.3%58.52026-09-15
56Qwen3 7maxAlibaba38.8%61.02026-09-15
57Nemotron 3 Ultra 550B A55BNVIDIAopen ↗38.6%48.12026-09-15
58Gemini 2.5 FlashGoogle38.4%52.52026-09-15
59Grok 4xAI38.1%58.52026-09-15
60Grok 4.3xAI38.1%50.52026-09-15

Cite as: BenchLeader, “MedCode leaderboard”, https://www.benchleader.com/benchmarks/vals_medcode, data as of 19 Sept 2026.

MedCode: questions

What does MedCode measure?
Clinical notes are given and the model must assign the correct medical billing codes, a task with strict, checkable answers. Scores are reported in percent of tasks solved; higher is better.
Which AI model leads MedCode?
Claude Opus 5 leads MedCode with 63.6% as of 19 Sept 2026, ahead of Gemini 3.1 Pro at 59.1%.
How many models have MedCode results?
90 model configurations have a MedCode result on BenchLeader, all taken from Vals AI.
Who runs MedCode and how often is it updated?
MedCode is published by Vals AI. BenchLeader re-reads the published results every morning and records the date each result was published.
Does MedCode count toward the BenchLeader Index?
No. MedCode is shown for reference but left out of the composite index.