BenchLeader

EnterpriseOps-Gym

Enterprise operations tasks in a simulated company: tickets, approvals and systems a model has to drive to finish a job.

As of 22 Sept 2026, Claude Fable 5 leads EnterpriseOps-Gym on BenchLeader with 51.1%, ahead of Gemini 3.7 Flash at 50.4%, across 24 model configurations with a published result.

Published by
Artificial Analysis
Category
Agents & tools
Index weight
Reference only
Models
24
Data as of
22 Sept 2026

Source: Artificial Analysis (artificialanalysis.ai). Data taken from the public leaderboard.

24 of 24
#
1Claude Fable 5thinkingAnthropic51.1%70.32026-06-09
2Gemini 3.7 FlashmediumGoogle50.4%64.42026-08-13
3Gemini 3.5 FlashhighGoogle50.1%63.42026-05-19
4DeepSeek V4 PromaxDeepSeekopen ↗49.6%63.72026-08-13
5Grok 4.6highSpaceXAI48.3%64.22026-08-12
6Claude Opus 5maxAnthropic47.5%69.52026-07-24
7Qwen3.8 2.4T A95BAlibabaopen ↗47.4%64.52026-08-12
8Muse Spark 1.2xhighMeta47.3%63.72026-08-05
9Kimi K3maxMoonshot AIopen ↗45.3%66.72026-07-16
10Claude Sonnet 5maxAnthropic44.7%60.02026-06-30
11Qwen3.8 27BxhighAlibabaopen ↗44.2%58.92026-08-14
12GPT-5.6 SolmaxOpenAI42.9%68.52026-07-09
13Gemini 3.5 Flash LiteGoogle42.4%55.22026-07-21
14GPT-5.6 LunamaxOpenAI40.8%59.92026-07-09
15GPT-5.6 TerramaxOpenAI38.5%64.82026-07-09
16GLM 5.3maxZhipu AIopen ↗36.4%65.42026-08-18
17Muse GlimmerhighMetaopen ↗34.7%52.52026-08-10
18Mistral Medium 3.5Mistral AIopen33.7%51.12026-04-29
19GLM 5.3 FlashZhipu AIopen ↗33.2%63.62026-08-26
20MiniMax M3MiniMaxopen ↗32.1%57.22026-06-01
21Nemotron 3 Ultra 550B A55BthinkingNVIDIAopen ↗28.9%57.32026-06-04
22gpt-oss-120bhighOpenAIopen ↗25.5%49.22025-08-05
23Nemotron 3 Super 120B A12bthinkingNVIDIAopen ↗23.6%51.22026-03-11
24Nemotron 3.5 LightningNVIDIAopen ↗18.4%48.02026-08-11

Cite as: BenchLeader, “EnterpriseOps-Gym leaderboard”, https://www.benchleader.com/benchmarks/aa_enterprise_ops_gym, data as of 22 Sept 2026.

EnterpriseOps-Gym: questions

What does EnterpriseOps-Gym measure?
Enterprise operations tasks in a simulated company: tickets, approvals and systems a model has to drive to finish a job. Scores are reported in percent of tasks solved; higher is better.
Which AI model leads EnterpriseOps-Gym?
Claude Fable 5 leads EnterpriseOps-Gym with 51.1% as of 22 Sept 2026, ahead of Gemini 3.7 Flash at 50.4%.
How many models have EnterpriseOps-Gym results?
24 model configurations have a EnterpriseOps-Gym result on BenchLeader, all taken from Artificial Analysis.
Who runs EnterpriseOps-Gym and how often is it updated?
EnterpriseOps-Gym is published by Artificial Analysis. BenchLeader re-reads the published results every morning and records the date each result was published.
Does EnterpriseOps-Gym count toward the BenchLeader Index?
No. EnterpriseOps-Gym is shown for reference but left out of the composite index.