Alibaba Open-Sources CommerceAgentBench for Long-Horizon Agent Testing
mhdfaran · x · 2026-08-27
The Alibaba International team has open-sourced CommerceAgentBench, a benchmark designed to evaluate long-horizon agents. Unlike standard Q&A benchmarks, it places an agent in a simulated environment of real online services—such as an inbox filled with 300 messy procurement emails—and tasks it with making the right buying decisions. This setup tests the agent's ability to handle complex, stateful, and long-duration business workflows. The project is now open source and welcomes contributions for mock environments.
More from coding & agent
- Agno Launches: A 'Self-Building' Agent Platform for the Cloud — pritisinghhhh · 2026-08-27
- Agent Opens Bambu Handy App on Phone to Reprint Job — haydendevs · 2026-08-27
- GitHub Copilot Teams update released with Slack integration — marlene_zw · 2026-08-27
- Connecting Mobile/Cloud Agents to Reach Local Beeper MCP — Basic-Let6828 · 2026-08-27
- Benchmarking DeepSeek V4 vs Qwen 3.8 on DGX Sparks — Legitimate_Hat_7852 · 2026-08-27
- Docker Is Not a Real Sandbox for Agent Code: From Containers to microVMs — aidenclarke_12 · 2026-08-27