New Leaderboard for Enterprise Operations Agents
hugo_larochelle · x · 2026-07-14
Artificial Analysis and ServiceNow have released EnterpriseOps-Gym-AA, an independent leaderboard revamped from the original EnterpriseOps-Gym. It evaluates whether AI agents can execute multi-step, stateful operational tasks within real enterprise systems while adhering to business rules and policies.
The benchmark runs via the Stirrup agent harness, covering 8 business domains. Rather than grading intermediate steps, scoring is based on the final state of the underlying database after the task. Currently, Anthropic's Claude Fable 5 (max) leads with 51%, followed by Google's Gemini 3.5 Flash (high) at 50%, and OpenAI's GPT-5.5 (xhigh) at 47%. Z.ai's GLM-5.2 (max) is the top-scoring open-weight model at 43%.
More from Research
- NeurIPS 2026 workshop will focus on on-device intelligence and local execution — YiMaTweets · 2026-07-21
- NeurIPS 2026 workshop calls papers on on-device intelligence — YiMaTweets · 2026-07-21
- AI Security Institute says every tested model tried to cheat in cyber evaluations — connoraxiotes · 2026-07-21
- AI companies are buying old books to avoid training on AI-generated slop — CackleRooster · 2026-07-21
- Sakana says multiple diffusion models plus MCTS beat test-time scaling on coding and math — SakanaAILabs · 2026-07-21
- Soofi S 30B-A3B releases a full pretraining report and claims open-model leads in English and German — abursuc · 2026-07-21