Alibaba's CommerceAgentBench: 107 Real E-Commerce Tasks, Top Model Fails 40%
iamfakhrealam · x · 2026-08-26
The Accio team at Alibaba released CommerceAgentBench, claiming it to be the hardest e-commerce agent benchmark.
- Data Source: Built on 27 years of real-world Alibaba e-commerce data, featuring 107 authentic business tasks.
- Task Nature: Designed to evaluate long-horizon business workflow execution, not just Q&A capabilities.
- Current Results: The leading model scores only 61.68% accuracy, failing to complete 41 real business tasks. This gap highlights the significant barrier for current agents in commercial deployment.
- Open Source: The team invites contributions for more mock environments and challenging tasks.
Related event: Alibaba Open-Sources CommerceAgentBench; Best Agent Completes Only 61.68%(7 posts)→
More from coding & agent
- AI Agent governance may become the next IAM security challenge — ingliguori · 2026-08-26
- Startup Runable demo: Agents can analyze business churn autonomously — HaktanSuren · 2026-08-26
- Goliath: Agent-first Headless Browser Automation Server with Persistent Firefox Sessions — ccerrato147 · 2026-08-26
- MemMiner: an MCP service giving AI agents shared memory via a rated knowledge graph — AdSpecialist4695 · 2026-08-26
- Archify: Agent Skill for Beautiful, Verifiable Architecture Diagrams — tt-a1i · 2026-08-26
- ConardLi Open Sources Skills Library for Web Design and RAG — ConardLi · 2026-08-26