BenchLeader

APEX-Agents (AA)

Artificial Analysis' run of Mercor's long-horizon professional agent tasks.

As of 19 Sept 2026, Gemini 3.5 Flash leads APEX-Agents (AA) on BenchLeader with 47.0%, ahead of Kimi K3 at 41.3%, across 31 model configurations with a published result.

Published by
Artificial Analysis
Category
Agents & tools
Index weight
Reference only
Models
31
Data as of
19 Sept 2026

Source: Artificial Analysis (artificialanalysis.ai). Data taken from the public leaderboard.

What the test looks like

APEX-Agents sets long-horizon professional tasks written and graded by domain experts at Mercor.

How it is scored

Percent of tasks judged successful, run by Artificial Analysis.

What to keep in mind

The Epoch-mirrored official results are already shown; this is an independent re-run.

31 of 31
#
1Gemini 3.5 FlashhighGoogle47.0%63.6
2Kimi K3maxMoonshot AIopen ↗41.3%67.2
3GPT-5.6 TerramaxOpenAI38.9%65.0
4GPT-5.5xhighOpenAI37.7%67.6
5GPT-5.6 LunamaxOpenAI35.8%60.2
6GLM-5.2maxZhipu AIopen ↗33.7%63.9
7GPT-5.4xhighOpenAI33.3%65.4
8Claude Opus 4.6maxAnthropic33.0%59.2
9Gemini 3.1 ProGoogle32.0%63.9
10Apodex 1.1Apodex31.2%56.1
11Kimi K2.6Moonshot AIopen ↗28.5%60.5
12GPT-5.4 minixhighOpenAI28.2%55.7
13Claude Sonnet 4.6maxAnthropic28.0%57.6
14Gemini 3 FlashthinkingGoogle27.7%60.4
15Ling 3.0 Flash FinUnknownopen27.4%55.4
16GPT-5.4 nanoxhighOpenAI24.9%54.6
17DeepSeek V4 PromaxDeepSeekopen ↗24.3%64.2
18Qwen3.7 PlusAlibaba22.4%59.9
19Grok 4.3highxAI17.0%58.2
20Qwen3.5 397B-A17BthinkingAlibabaopen ↗15.3%56.1
21Step 3.7 FlashStepFunopen ↗14.8%52.6
22DeepSeek V3.2thinkingDeepSeekopen14.5%56.1
23GLM-5thinkingZhipu AIopen14.4%59.6
24Grok 4.20thinkingxAI14.2%59.9
25Gemini 3.1 Flash LiteGoogle12.2%54.5
26Kimi K2.5thinkingMoonshot AIopen11.5%58.5
27MiniMax-M2.7MiniMaxopen ↗10.6%56.1
28gpt-oss-120bhighOpenAIopen3.1%49.7
29MiMo-V2.5-ProXiaomiopen ↗2.4%59.2
30Nemotron 3 SuperthinkingNVIDIAopen ↗1.8%51.5
31gpt-oss-20bhighOpenAIopen0.7%44.7

Cite as: BenchLeader, “APEX-Agents (AA) leaderboard”, https://www.benchleader.com/benchmarks/aa_apex_agents, data as of 19 Sept 2026.

APEX-Agents (AA): questions

What does APEX-Agents (AA) measure?
APEX-Agents sets long-horizon professional tasks written and graded by domain experts at Mercor. Scores are reported in percent of tasks solved; higher is better.
Which AI model leads APEX-Agents (AA)?
Gemini 3.5 Flash leads APEX-Agents (AA) with 47.0% as of 19 Sept 2026, ahead of Kimi K3 at 41.3%.
How many models have APEX-Agents (AA) results?
31 model configurations have a APEX-Agents (AA) result on BenchLeader, all taken from Artificial Analysis.
Who runs APEX-Agents (AA) and how often is it updated?
APEX-Agents (AA) is published by Artificial Analysis. BenchLeader re-reads the published results every morning and records the date each result was published.
Does APEX-Agents (AA) count toward the BenchLeader Index?
No. APEX-Agents (AA) is shown for reference but left out of the composite index.