EnterpriseOps-Gym
Enterprise operations tasks in a simulated company: tickets, approvals and systems a model has to drive to finish a job.
As of 22 Sept 2026, Claude Fable 5 leads EnterpriseOps-Gym on BenchLeader with 51.1%, ahead of Gemini 3.7 Flash at 50.4%, across 24 model configurations with a published result.
- Published by
- Artificial Analysis
- Category
- Agents & tools
- Index weight
- Reference only
- Models
- 24
- Data as of
- 22 Sept 2026
Source: Artificial Analysis (artificialanalysis.ai). Data taken from the public leaderboard.
- 1Claude Fable 5 (thinking)51.1%
- 2Gemini 3.7 Flash (medium)50.4%
- 3Gemini 3.5 Flash (high)50.1%
- 4DeepSeek V4 Pro (max)49.6%
- 5Grok 4.6 (high)48.3%
- 6Claude Opus 5 (max)47.5%
- 7Qwen3.8 2.4T A95B47.4%
- 8Muse Spark 1.2 (xhigh)47.3%
- 9Kimi K3 (max)45.3%
- 10Claude Sonnet 5 (max)44.7%
- 11Qwen3.8 27B (xhigh)44.2%
- 12GPT-5.6 Sol (max)42.9%
- 13Gemini 3.5 Flash Lite42.4%
- 14GPT-5.6 Luna (max)40.8%
- 15GPT-5.6 Terra (max)38.5%
24 of 24
| # | ||||
|---|---|---|---|---|
| 1 | 51.1% | 70.3 | 2026-06-09 | |
| 2 | 50.4% | 64.4 | 2026-08-13 | |
| 3 | 50.1% | 63.4 | 2026-05-19 | |
| 4 | 49.6% | 63.7 | 2026-08-13 | |
| 5 | 48.3% | 64.2 | 2026-08-12 | |
| 6 | 47.5% | 69.5 | 2026-07-24 | |
| 7 | 47.4% | 64.5 | 2026-08-12 | |
| 8 | 47.3% | 63.7 | 2026-08-05 | |
| 9 | 45.3% | 66.7 | 2026-07-16 | |
| 10 | 44.7% | 60.0 | 2026-06-30 | |
| 11 | 44.2% | 58.9 | 2026-08-14 | |
| 12 | 42.9% | 68.5 | 2026-07-09 | |
| 13 | 42.4% | 55.2 | 2026-07-21 | |
| 14 | 40.8% | 59.9 | 2026-07-09 | |
| 15 | 38.5% | 64.8 | 2026-07-09 | |
| 16 | 36.4% | 65.4 | 2026-08-18 | |
| 17 | 34.7% | 52.5 | 2026-08-10 | |
| 18 | 33.7% | 51.1 | 2026-04-29 | |
| 19 | 33.2% | 63.6 | 2026-08-26 | |
| 20 | 32.1% | 57.2 | 2026-06-01 | |
| 21 | 28.9% | 57.3 | 2026-06-04 | |
| 22 | 25.5% | 49.2 | 2025-08-05 | |
| 23 | 23.6% | 51.2 | 2026-03-11 | |
| 24 | 18.4% | 48.0 | 2026-08-11 |
Cite as: BenchLeader, “EnterpriseOps-Gym leaderboard”, https://www.benchleader.com/benchmarks/aa_enterprise_ops_gym, data as of 22 Sept 2026.
EnterpriseOps-Gym: questions
- What does EnterpriseOps-Gym measure?
- Enterprise operations tasks in a simulated company: tickets, approvals and systems a model has to drive to finish a job. Scores are reported in percent of tasks solved; higher is better.
- Which AI model leads EnterpriseOps-Gym?
- Claude Fable 5 leads EnterpriseOps-Gym with 51.1% as of 22 Sept 2026, ahead of Gemini 3.7 Flash at 50.4%.
- How many models have EnterpriseOps-Gym results?
- 24 model configurations have a EnterpriseOps-Gym result on BenchLeader, all taken from Artificial Analysis.
- Who runs EnterpriseOps-Gym and how often is it updated?
- EnterpriseOps-Gym is published by Artificial Analysis. BenchLeader re-reads the published results every morning and records the date each result was published.
- Does EnterpriseOps-Gym count toward the BenchLeader Index?
- No. EnterpriseOps-Gym is shown for reference but left out of the composite index.