BenchLeader

Analyst Agent (AA)

Multi-step financial-analyst tasks with tools, run by Artificial Analysis.

As of 19 Sept 2026, Gemini 3.7 Flash leads Analyst Agent (AA) on BenchLeader with 60.0%, ahead of Claude Fable 5.1 at 57.5%, across 32 model configurations with a published result.

Published by
Artificial Analysis
Category
Agents & tools
Index weight
Reference only
Models
32
Data as of
19 Sept 2026

Source: Artificial Analysis (artificialanalysis.ai). Data taken from the public leaderboard.

What the test looks like

The agent is given an analyst brief, a set of documents and tools, and must produce the analysis a junior analyst would.

How it is scored

Rubric score across tasks, run by Artificial Analysis.

What to keep in mind

New, with a small number of models measured so far.

32 of 32
#
1Gemini 3.7 FlashhighGoogle60.0%64.6
2Claude Fable 5.1thinkingAnthropic57.5%71.3
3Claude Opus 5maxAnthropic53.8%69.9
4GPT-6 AstramaxOpenAI51.3%71.8
5GPT-5.5xhighOpenAI50.0%67.6
6Claude Fable 5thinkingAnthropic48.8%70.9
7GPT-5.6 SolmaxOpenAI47.5%68.8
8Claude Sonnet 5maxAnthropic46.3%60.5
9Claude Opus 4.8maxAnthropic45.0%64.2
10Gemini 3.5 FlashhighGoogle45.0%63.6
11Claude Opus 4.7maxAnthropic43.8%64.3
12Grok 4.6highxAI41.3%64.5
13Gemini 3.1 ProGoogle41.3%63.9
14Kimi K3maxMoonshot AIopen ↗38.8%67.2
15Grok 4.5highxAI35.0%61.0
16Inkling SmallThinking Machinesopen ↗27.5%56.3
17DeepSeek V4 FlashmaxDeepSeekopen ↗25.0%62.1
18MiMo-V2.5-ProXiaomiopen ↗20.0%59.2
19Claude Sonnet 4.6maxAnthropic20.0%57.6
20DeepSeek V4 PromaxDeepSeekopen ↗18.8%64.2
21Qwen3 7maxAlibaba18.8%61.0
22Ling 3.0 Flash FinUnknownopen16.3%55.4
23GPT-5.4 minixhighOpenAI15.0%55.7
24Claude Haiku 4.5thinkingAnthropic15.0%47.5
25Mistral Medium 3.5Mistral AIopen12.5%51.3
26MiniMax-M2.7MiniMaxopen ↗11.3%56.1
27MiniMax-M3MiniMaxopen ↗10.0%57.5
28Grok 4.3highxAI8.8%58.2
29Gemini 3.1 Flash LiteGoogle8.8%54.5
30GPT-5.5no reasoningOpenAI7.5%54.0
31Nemotron 3 Ultra 550B A55BthinkingNVIDIAopen ↗6.3%57.6
32Mistral Small 4thinkingMistral AIopen1.3%46.3

Cite as: BenchLeader, “Analyst Agent (AA) leaderboard”, https://www.benchleader.com/benchmarks/aa_analyst_agent, data as of 19 Sept 2026.

Analyst Agent (AA): questions

What does Analyst Agent (AA) measure?
The agent is given an analyst brief, a set of documents and tools, and must produce the analysis a junior analyst would. Scores are reported in percent of tasks solved; higher is better.
Which AI model leads Analyst Agent (AA)?
Gemini 3.7 Flash leads Analyst Agent (AA) with 60.0% as of 19 Sept 2026, ahead of Claude Fable 5.1 at 57.5%.
How many models have Analyst Agent (AA) results?
32 model configurations have a Analyst Agent (AA) result on BenchLeader, all taken from Artificial Analysis.
Who runs Analyst Agent (AA) and how often is it updated?
Analyst Agent (AA) is published by Artificial Analysis. BenchLeader re-reads the published results every morning and records the date each result was published.
Does Analyst Agent (AA) count toward the BenchLeader Index?
No. Analyst Agent (AA) is shown for reference but left out of the composite index.