Alibaba Open-Sources CommerceAgentBench for Long-Horizon Agent Testing

mhdfaran · x · 2026-08-27

The Alibaba International team has open-sourced CommerceAgentBench, a benchmark designed to evaluate long-horizon agents. Unlike standard Q&A benchmarks, it places an agent in a simulated environment of real online services—such as an inbox filled with 300 messy procurement emails—and tasks it with making the right buying decisions. This setup tests the agent's ability to handle complex, stateful, and long-duration business workflows. The project is now open source and welcomes contributions for mock environments.

Related event: Alibaba Open-Sources CommerceAgentBench, Toughest E-commerce Agent Benchmark(11 posts)→

Original post →

More from coding & agent

coding & agent channel →