Alibaba Open-Sources CommerceAgentBench: Best Agent Completes Only 61.68% of Real E-commerce Tasks

Alibaba's international marketplace team has released and open-sourced CommerceAgentBench (GitHub repo Accio-org/CommerceAgentBench, already at 1.2k stars), a benchmark for testing whether AI Agents can actually get real e-commerce work done rather than just give correct answers. The benchmark contains 107 tasks covering the full chain—procurement, product listing, operations, fulfillment, and after-sales—distilled from real Alibaba international e-commerce data, with sources including 10 million SMB users and 1.6 million complete conversations. The best-performing agent currently completes only 61.68% of tasks.

Confirmed

Why it matters

2026-08-26 ~ 2026-08-26 · 5 related posts

Primary sources