APEX-Agents (AA)
Artificial Analysis' run of Mercor's long-horizon professional agent tasks.
As of 19 Sept 2026, Gemini 3.5 Flash leads APEX-Agents (AA) on BenchLeader with 47.0%, ahead of Kimi K3 at 41.3%, across 31 model configurations with a published result.
- Published by
- Artificial Analysis
- Category
- Agents & tools
- Index weight
- Reference only
- Models
- 31
- Data as of
- 19 Sept 2026
Source: Artificial Analysis (artificialanalysis.ai). Data taken from the public leaderboard.
What the test looks like
APEX-Agents sets long-horizon professional tasks written and graded by domain experts at Mercor.
How it is scored
Percent of tasks judged successful, run by Artificial Analysis.
What to keep in mind
The Epoch-mirrored official results are already shown; this is an independent re-run.
- 1Gemini 3.5 Flash (high)47.0%
- 2Kimi K3 (max)41.3%
- 3GPT-5.6 Terra (max)38.9%
- 4GPT-5.5 (xhigh)37.7%
- 5GPT-5.6 Luna (max)35.8%
- 6GLM-5.2 (max)33.7%
- 7GPT-5.4 (xhigh)33.3%
- 8Claude Opus 4.6 (max)33.0%
- 9Gemini 3.1 Pro32.0%
- 10Apodex 1.131.2%
- 11Kimi K2.628.5%
- 12GPT-5.4 mini (xhigh)28.2%
- 13Claude Sonnet 4.6 (max)28.0%
- 14Gemini 3 Flash (thinking)27.7%
- 15Ling 3.0 Flash Fin27.4%
31 of 31
| # | |||
|---|---|---|---|
| 1 | 47.0% | 63.6 | |
| 2 | 41.3% | 67.2 | |
| 3 | 38.9% | 65.0 | |
| 4 | 37.7% | 67.6 | |
| 5 | 35.8% | 60.2 | |
| 6 | 33.7% | 63.9 | |
| 7 | 33.3% | 65.4 | |
| 8 | 33.0% | 59.2 | |
| 9 | 32.0% | 63.9 | |
| 10 | Apodex 1.1Apodex | 31.2% | 56.1 |
| 11 | 28.5% | 60.5 | |
| 12 | 28.2% | 55.7 | |
| 13 | 28.0% | 57.6 | |
| 14 | 27.7% | 60.4 | |
| 15 | Ling 3.0 Flash FinUnknownopen | 27.4% | 55.4 |
| 16 | 24.9% | 54.6 | |
| 17 | 24.3% | 64.2 | |
| 18 | 22.4% | 59.9 | |
| 19 | 17.0% | 58.2 | |
| 20 | 15.3% | 56.1 | |
| 21 | 14.8% | 52.6 | |
| 22 | 14.5% | 56.1 | |
| 23 | 14.4% | 59.6 | |
| 24 | 14.2% | 59.9 | |
| 25 | 12.2% | 54.5 | |
| 26 | 11.5% | 58.5 | |
| 27 | 10.6% | 56.1 | |
| 28 | 3.1% | 49.7 | |
| 29 | 2.4% | 59.2 | |
| 30 | 1.8% | 51.5 | |
| 31 | 0.7% | 44.7 |
Cite as: BenchLeader, “APEX-Agents (AA) leaderboard”, https://www.benchleader.com/benchmarks/aa_apex_agents, data as of 19 Sept 2026.
APEX-Agents (AA): questions
- What does APEX-Agents (AA) measure?
- APEX-Agents sets long-horizon professional tasks written and graded by domain experts at Mercor. Scores are reported in percent of tasks solved; higher is better.
- Which AI model leads APEX-Agents (AA)?
- Gemini 3.5 Flash leads APEX-Agents (AA) with 47.0% as of 19 Sept 2026, ahead of Kimi K3 at 41.3%.
- How many models have APEX-Agents (AA) results?
- 31 model configurations have a APEX-Agents (AA) result on BenchLeader, all taken from Artificial Analysis.
- Who runs APEX-Agents (AA) and how often is it updated?
- APEX-Agents (AA) is published by Artificial Analysis. BenchLeader re-reads the published results every morning and records the date each result was published.
- Does APEX-Agents (AA) count toward the BenchLeader Index?
- No. APEX-Agents (AA) is shown for reference but left out of the composite index.