BenchLeader

Excel Modeling Benchmark

Building financial models in Excel from a brief. Run by Vals AI.

As of 19 Sept 2026, Claude Fable 5.1 leads Excel Modeling Benchmark on BenchLeader with 76.7%, ahead of Claude Fable 5 at 73.7%, across 55 model configurations with a published result.

Published by
Vals AI
Category
Knowledge
Index weight
Reference only
Models
55
Data as of
19 Sept 2026

Vals AI (vals.ai).

What the test looks like

The model must produce a working Excel financial model, with formulas and structure, from an analyst-style brief.

How it is scored

Rubric score across models built, run by Vals AI.

What to keep in mind

Tool access and file handling matter as much as finance knowledge.

55 of 55
#
1Claude Fable 5.1Anthropic76.7%64.52026-09-15
2Claude Fable 5Anthropic73.7%68.32026-09-15
3Claude Opus 5Anthropic73.6%66.72026-09-15
4GPT-5.6 SolmaxOpenAI72.3%68.82026-09-15
5Gemini 3.8 FlashhighGoogle72.2%64.52026-09-15
6GPT-6 AstramaxOpenAI71.7%71.82026-09-15
7Gemini 3.7 FlashhighGoogle71.3%64.62026-09-15
8Claude Opus 4.8Anthropic69.4%61.92026-09-15
9Muse Spark 1.3maxMeta67.4%69.32026-09-15
10GPT-5.6 LunamaxOpenAI67.1%60.22026-09-15
11Kimi K3maxMoonshot AIopen ↗66.4%67.22026-09-15
12Claude Sonnet 5Anthropic66.3%57.22026-09-15
13GPT-5.6 TerramaxOpenAI66.2%65.02026-09-15
14Gemini 3.6 FlashhighGoogle65.4%61.72026-09-15
15GPT-5.5xhighOpenAI64.5%67.62026-09-15
16Claude Opus 4.7Anthropic63.7%64.52026-09-15
17Gemini 3.5 FlashhighGoogle63.5%63.62026-09-15
18Grok 4.6highxAI62.7%64.52026-09-15
19Muse Spark 1.3xhighMeta62.7%68.02026-09-15
20GLM-5.2Zhipu AIopen ↗61.5%52.12026-09-15
21Claude Sonnet 4.6Anthropic60.1%59.02026-09-15
22Qwen3 8maxAlibaba60.1%66.52026-09-15
23Qwen3.8 27BxhighAlibabaopen ↗59.7%59.22026-09-15
24Kimi K2.6Moonshot AIopen ↗57.9%60.52026-09-15
25DeepSeek V4.1 FlashhighDeepSeekopen ↗57.2%2026-09-15
26Qwen3 7maxAlibaba57.0%61.02026-09-15
27Muse Spark 1.2xhighMeta57.0%64.12026-09-15
28DeepSeek V4 Flash 0731highDeepSeekopen ↗57.0%55.02026-09-15
29Muse Spark 1.1xhighMeta56.4%61.82026-09-15
30GLM-5.3maxZhipu AIopen ↗56.3%65.82026-09-15
31GLM-5.3-FlashmaxZhipu AIopen ↗55.9%55.32026-09-15
32MiMo-V2.5-ProXiaomiopen ↗55.2%59.22026-09-15
33Grok 4.5highxAI52.9%61.02026-09-15
34DeepSeek V4 Pro 0813maxDeepSeekopen ↗52.8%56.02026-09-15
35Gemini 3.1 ProhighGoogle52.6%60.72026-09-15
36DeepSeek V4 PromaxDeepSeekopen ↗51.6%64.22026-09-15
37Qwen3.7 PlusAlibaba49.4%59.92026-09-15
38MiniMax-M3MiniMaxopen ↗47.8%57.52026-09-15
39GPT-5.4 minixhighOpenAI45.4%55.72026-09-15
40GPT-5.4 nanohighOpenAI44.8%49.12026-09-15
41Gemini 3.5 Flash LitehighGoogle43.2%44.32026-09-15
42Qwen3.6 PlusAlibabaopen32.9%58.52026-09-15
43Inkling SmallThinking Machinesopen ↗31.8%56.32026-09-15
44Nemotron 3 Ultra 550B A55BNVIDIAopen ↗31.7%48.12026-09-15
45Kimi K2.5thinkingMoonshot AIopen28.5%58.52026-09-15
46MiniMax-M2.7MiniMaxopen ↗28.4%56.12026-09-15
47Gemini 3 FlashhighGoogle26.0%58.72026-09-15
48Ling 3.0 FlashAnt Groupopen ↗25.0%53.62026-09-15
49Claude Haiku 4.5thinkingAnthropic22.7%47.52026-09-15
50Grok 4.3highxAI18.4%58.22026-09-15
51Nemotron Lightning 3.5 30B A3bNVIDIA17.0%2026-09-15
52Grok 4.20thinkingxAI11.7%59.92026-09-15
53Mistral Medium 3.5highMistral AIopen11.2%2026-09-15
54Mercury 2.5highInception8.8%2026-09-15
55Gemini 3.1 Flash LitehighGoogle8.6%49.12026-09-15

Cite as: BenchLeader, “Excel Modeling Benchmark leaderboard”, https://www.benchleader.com/benchmarks/vals_emb, data as of 19 Sept 2026.

Excel Modeling Benchmark: questions

What does Excel Modeling Benchmark measure?
The model must produce a working Excel financial model, with formulas and structure, from an analyst-style brief. Scores are reported in percent of tasks solved; higher is better.
Which AI model leads Excel Modeling Benchmark?
Claude Fable 5.1 leads Excel Modeling Benchmark with 76.7% as of 19 Sept 2026, ahead of Claude Fable 5 at 73.7%.
How many models have Excel Modeling Benchmark results?
55 model configurations have a Excel Modeling Benchmark result on BenchLeader, all taken from Vals AI.
Who runs Excel Modeling Benchmark and how often is it updated?
Excel Modeling Benchmark is published by Vals AI. BenchLeader re-reads the published results every morning and records the date each result was published.
Does Excel Modeling Benchmark count toward the BenchLeader Index?
No. Excel Modeling Benchmark is shown for reference but left out of the composite index.