SAGE
Grading handwritten maths work from images. Run by Vals AI.
As of 19 Sept 2026, Claude Opus 4.7 leads SAGE on BenchLeader with 56.1%, ahead of Gemma 4 31B at 55.0%, across 79 model configurations with a published result.
- Published by
- Vals AI
- Category
- Multimodal
- Index weight
- Reference only
- Models
- 79
- Data as of
- 19 Sept 2026
Vals AI (vals.ai).
What the test looks like
Images of handwritten student maths work must be graded against a rubric, requiring reading handwriting and checking reasoning.
How it is scored
Agreement with expert graders, run by Vals AI.
What to keep in mind
Handwriting quality and image resolution affect results.
- 1Claude Opus 4.756.1%
- 2Gemma 4 31B (high)55.0%
- 3Claude Opus 4.854.8%
- 4Kimi K354.3%
- 5GPT-5.6 Sol (max)52.6%
- 6Qwen3.8 27B (xhigh)52.4%
- 7Claude Opus 4.5 (thinking)52.1%
- 8Claude Fable 551.9%
- 9Gemini 3 Flash (high)51.9%
- 10Claude Opus 4.6 (thinking)51.6%
- 11GPT-5.5 (xhigh)51.5%
- 12Qwen3 8 (max)51.3%
- 13GPT-5.4 mini (xhigh)50.8%
- 14MiniMax-M350.6%
- 15Kimi K2.650.2%
79 of 79
| # | ||||
|---|---|---|---|---|
| 1 | 56.1% | 64.5 | 2026-09-10 | |
| 2 | 55.0% | – | 2026-09-10 | |
| 3 | 54.8% | 61.9 | 2026-09-10 | |
| 4 | 54.3% | 64.2 | 2026-09-10 | |
| 5 | 52.6% | 68.8 | 2026-09-10 | |
| 6 | 52.4% | 59.2 | 2026-09-10 | |
| 7 | 52.1% | 60.8 | 2026-09-10 | |
| 8 | 51.9% | 68.3 | 2026-09-10 | |
| 9 | 51.9% | 58.7 | 2026-09-10 | |
| 10 | 51.6% | 62.5 | 2026-09-10 | |
| 11 | 51.5% | 67.6 | 2026-09-10 | |
| 12 | 51.3% | 66.5 | 2026-09-10 | |
| 13 | 50.8% | 55.7 | 2026-09-10 | |
| 14 | 50.6% | 57.5 | 2026-09-10 | |
| 15 | 50.2% | 60.5 | 2026-09-10 | |
| 16 | 49.9% | 63.6 | 2026-09-10 | |
| 17 | 49.9% | 58.5 | 2026-09-10 | |
| 18 | 49.5% | 49.1 | 2026-09-10 | |
| 19 | 49.4% | 66.7 | 2026-09-10 | |
| 20 | 49.3% | 62.0 | 2026-09-10 | |
| 21 | 49.2% | 64.6 | 2026-09-10 | |
| 22 | 48.9% | 57.2 | 2026-09-10 | |
| 23 | 48.7% | 61.7 | 2026-09-10 | |
| 24 | 48.7% | 60.7 | 2026-09-10 | |
| 25 | 48.5% | 64.5 | 2026-09-10 | |
| 26 | 47.9% | – | 2026-09-10 | |
| 27 | 47.7% | 64.1 | 2026-09-10 | |
| 28 | 47.6% | 61.1 | 2026-09-10 | |
| 29 | 47.3% | 44.3 | 2026-09-10 | |
| 30 | 47.0% | 65.0 | 2026-09-10 | |
| 31 | 46.6% | 59.0 | 2026-09-10 | |
| 32 | 46.4% | 71.8 | 2026-09-10 | |
| 33 | 45.6% | 61.8 | 2026-09-10 | |
| 34 | 45.6% | 44.7 | 2026-09-10 | |
| 35 | 45.0% | 58.2 | 2026-09-10 | |
| 36 | 44.9% | 58.5 | 2026-09-10 | |
| 37 | 44.8% | 52.5 | 2026-09-10 | |
| 38 | 44.2% | 60.2 | 2026-09-10 | |
| 39 | 44.1% | 52.1 | 2026-09-10 | |
| 40 | 43.7% | 58.6 | 2026-09-10 | |
| 41 | 43.3% | 65.4 | 2026-09-10 | |
| 42 | 43.2% | 58.7 | 2026-09-10 | |
| 43 | 43.0% | 55.3 | 2026-09-10 | |
| 44 | 42.5% | 52.5 | 2026-09-10 | |
| 45 | 42.1% | – | 2026-09-10 | |
| 46 | 41.9% | 54.5 | 2026-09-10 | |
| 47 | 41.8% | 53.0 | 2026-09-10 | |
| 48 | 41.1% | 50.8 | 2026-09-10 | |
| 49 | 39.6% | 48.4 | 2026-09-10 | |
| 50 | 39.4% | 51.6 | 2026-09-10 | |
| 51 | 39.3% | 59.9 | 2026-09-10 | |
| 52 | 38.2% | 59.9 | 2026-09-10 | |
| 53 | 38.1% | 49.1 | 2026-09-10 | |
| 54 | 37.6% | – | 2026-09-10 | |
| 55 | 36.1% | 53.4 | 2026-09-10 | |
| 56 | 35.7% | 39.4 | 2026-09-10 | |
| 57 | 35.1% | 64.5 | 2026-09-10 | |
| 58 | 35.0% | 49.7 | 2026-09-10 | |
| 59 | 35.0% | 61.0 | 2026-09-10 | |
| 60 | 34.8% | 36.4 | 2026-09-10 |
Cite as: BenchLeader, “SAGE leaderboard”, https://www.benchleader.com/benchmarks/vals_sage, data as of 19 Sept 2026.
SAGE: questions
- What does SAGE measure?
- Images of handwritten student maths work must be graded against a rubric, requiring reading handwriting and checking reasoning. Scores are reported in percent of tasks solved; higher is better.
- Which AI model leads SAGE?
- Claude Opus 4.7 leads SAGE with 56.1% as of 19 Sept 2026, ahead of Gemma 4 31B at 55.0%.
- How many models have SAGE results?
- 79 model configurations have a SAGE result on BenchLeader, all taken from Vals AI.
- Who runs SAGE and how often is it updated?
- SAGE is published by Vals AI. BenchLeader re-reads the published results every morning and records the date each result was published.
- Does SAGE count toward the BenchLeader Index?
- No. SAGE is shown for reference but left out of the composite index.