BenchLeader

SAGE

Grading handwritten maths work from images. Run by Vals AI.

As of 19 Sept 2026, Claude Opus 4.7 leads SAGE on BenchLeader with 56.1%, ahead of Gemma 4 31B at 55.0%, across 79 model configurations with a published result.

Published by
Vals AI
Category
Multimodal
Index weight
Reference only
Models
79
Data as of
19 Sept 2026

Vals AI (vals.ai).

What the test looks like

Images of handwritten student maths work must be graded against a rubric, requiring reading handwriting and checking reasoning.

How it is scored

Agreement with expert graders, run by Vals AI.

What to keep in mind

Handwriting quality and image resolution affect results.

79 of 79
#
1Claude Opus 4.7Anthropic56.1%64.52026-09-10
2Gemma 4 31BhighGoogleopen ↗55.0%2026-09-10
3Claude Opus 4.8Anthropic54.8%61.92026-09-10
4Kimi K3Moonshot AIopen ↗54.3%64.22026-09-10
5GPT-5.6 SolmaxOpenAI52.6%68.82026-09-10
6Qwen3.8 27BxhighAlibabaopen ↗52.4%59.22026-09-10
7Claude Opus 4.5thinkingAnthropic52.1%60.82026-09-10
8Claude Fable 5Anthropic51.9%68.32026-09-10
9Gemini 3 FlashhighGoogle51.9%58.72026-09-10
10Claude Opus 4.6thinkingAnthropic51.6%62.52026-09-10
11GPT-5.5xhighOpenAI51.5%67.62026-09-10
12Qwen3 8maxAlibaba51.3%66.52026-09-10
13GPT-5.4 minixhighOpenAI50.8%55.72026-09-10
14MiniMax-M3MiniMaxopen ↗50.6%57.52026-09-10
15Kimi K2.6Moonshot AIopen ↗50.2%60.52026-09-10
16Gemini 3.5 FlashhighGoogle49.9%63.62026-09-10
17Kimi K2.5thinkingMoonshot AIopen49.9%58.52026-09-10
18Gemini 3.1 Flash LitehighGoogle49.5%49.12026-09-10
19Claude Opus 5Anthropic49.4%66.72026-09-10
20GPT-5.2xhighOpenAI49.3%62.02026-09-10
21Gemini 3.7 FlashhighGoogle49.2%64.62026-09-10
22Claude Sonnet 5Anthropic48.9%57.22026-09-10
23Gemini 3.6 FlashhighGoogle48.7%61.72026-09-10
24Gemini 3.1 ProhighGoogle48.7%60.72026-09-10
25Claude Fable 5.1Anthropic48.5%64.52026-09-10
26DeepSeek V4.1 FlashhighDeepSeekopen ↗47.9%2026-09-10
27Muse Spark 1.2xhighMeta47.7%64.12026-09-10
28Gemini 3 ProhighGoogle47.6%61.12026-09-10
29Gemini 3.5 Flash LitehighGoogle47.3%44.32026-09-10
30GPT-5.6 TerramaxOpenAI47.0%65.02026-09-10
31Claude Sonnet 4.6Anthropic46.6%59.02026-09-10
32GPT-6 AstramaxOpenAI46.4%71.82026-09-10
33Muse Spark 1.1xhighMeta45.6%61.82026-09-10
34Qwen3.6 27BAlibaba45.6%44.72026-09-10
35Claude Opus 4.5Anthropic45.0%58.22026-09-10
36Qwen3.6 PlusAlibabaopen44.9%58.52026-09-10
37Gemini 2.5 FlashGoogle44.8%52.52026-09-10
38GPT-5.6 LunamaxOpenAI44.2%60.22026-09-10
39Gemini 2.5 Flash 09 2025Google44.1%52.12026-09-10
40GPT-5highOpenAI43.7%58.62026-09-10
41GPT-5.4xhighOpenAI43.3%65.42026-09-10
42GPT-5.1highOpenAI43.2%58.72026-09-10
43GPT-5 minihighOpenAI43.0%55.32026-09-10
44Qwen3.5 FlashAlibaba42.5%52.52026-09-10
45Qwen3-VL PlusAlibaba42.1%2026-09-10
46Gemini 2.5 ProGoogle41.9%54.52026-09-10
47o3highOpenAI41.8%53.02026-09-10
48o4-minihighOpenAI41.1%50.82026-09-10
49Gemini 2.5 FlashthinkingGoogle39.6%48.42026-09-10
50Gemini 2.5 Flash 09 2025thinkingGoogle39.4%51.62026-09-10
51Qwen3.7 PlusAlibaba39.3%59.92026-09-10
52Grok 4.20thinkingxAI38.2%59.92026-09-10
53GPT-5.4 nanohighOpenAI38.1%49.12026-09-10
54Mistral Medium 3.5highMistral AIopen37.6%2026-09-10
55Claude Sonnet 4.5thinkingAnthropic36.1%53.42026-09-10
56Llama 4 Maverick BasicMeta35.7%39.42026-09-10
57Gemini 3.8 FlashhighGoogle35.1%64.52026-09-10
58Claude Sonnet 4Anthropic35.0%49.72026-09-10
59Grok 4.5highxAI35.0%61.02026-09-10
60Llama Llama 4 Scout 17B 16eMeta34.8%36.42026-09-10

Cite as: BenchLeader, “SAGE leaderboard”, https://www.benchleader.com/benchmarks/vals_sage, data as of 19 Sept 2026.

SAGE: questions

What does SAGE measure?
Images of handwritten student maths work must be graded against a rubric, requiring reading handwriting and checking reasoning. Scores are reported in percent of tasks solved; higher is better.
Which AI model leads SAGE?
Claude Opus 4.7 leads SAGE with 56.1% as of 19 Sept 2026, ahead of Gemma 4 31B at 55.0%.
How many models have SAGE results?
79 model configurations have a SAGE result on BenchLeader, all taken from Vals AI.
Who runs SAGE and how often is it updated?
SAGE is published by Vals AI. BenchLeader re-reads the published results every morning and records the date each result was published.
Does SAGE count toward the BenchLeader Index?
No. SAGE is shown for reference but left out of the composite index.