Best Agent Completes Only 61.68%: Alibaba's CommerceAgentBench Goes Live on GitHub

alifcoder · x · 2026-08-26

CommerceAgentBench is now open-source on GitHub (Accio-org/CommerceAgentBench, 1.2k stars). Its 107 tasks were distilled from Alibaba's real e-commerce data: 10M SME users → 1.6M full-length conversations → 200K commercial execution trajectories → 2,000 high-value workflows → 107 reproducible tasks, drawing on 27 years of Alibaba e-commerce experience. The leaderboard shows the best observed result completed just 66 of 107 tasks (61.68%) — a model can sound convincing in chat yet break when the job demands messy inputs, conflicting information, business rules, multiple tools, real execution, and a verifiable final state. The team invites contributors to expand its mock environments. The author's takeaway: agent benchmarks should ask less "what did the AI say?" and more "what did the AI actually do?"

Related event: Alibaba Open-Sources CommerceAgentBench for Real E-commerce Tasks(3 posts)→

Original post →

More from coding & agent

coding & agent channel →