ITSMBench Released: Frontier Models Struggle with Enterprise Agent Reliability
Shahules786 · x · 2026-07-30
Vibrant Labs has open-sourced Enterprise Worlds, a suite of executable environments for training and evaluating AI agents on realistic enterprise workflows. The first release, ITSMBench, is an IT service management (ITSM) benchmark.
Core Design & Findings
- Environment: ITSMBench places agents inside a live ITSM environment featuring persistent enterprise state, simulated users, policy-governed workflows, and 93 typed tools.
- Evaluation: Tasks cover tickets, assets, approvals, and incident management. Success requires leaving the environment in the correct final state, not just producing a plausible transcript.
- Model Gaps: Results show that while frontier models can often solve a task once, their repeated-success scores drop sharply. They struggle significantly with policy-following, ambiguity resolution, and maintaining correct state across workflows.
- Failure Modes: Models often complete visible steps while missing enterprise consequences, such as reporting an SLA pause without updating the SLA row, inventing audit timestamps, skipping required questions, or sending premature customer emails.
Related event: Vibrant Labs Launches ITSMBench for Enterprise AI Agents(2 posts)→
More from coding & agent
- Self-Improving Agents Boost vLLM Inference Throughput by 16% for Trillion-Param Models — yisongyue · 2026-07-30
- Verdent Integrates Kimi K3 with Optimized Harness for Agentic Coding — eyishazyer · 2026-07-30
- MCP Drives Analytics Shift: Amplitude Says Half of Queries Will Be AI-Run — TansuYegen · 2026-07-30
- Verdent Partners with Moonshot to Deeply Optimize Kimi K3 for Agentic Coding — eyishazyer · 2026-07-30
- A Buyer's Guide to AI Agents: Three Questions to Ask Before Automating — AlexKim · 2026-07-30
- Armature Launches Product Analytics Tailored for MCPs and AI Apps — ycombinator · 2026-07-30