Alibaba Open-Sources CommerceAgentBench: Best Agent Completes Only 61.68% of Real E-commerce Tasks
Alibaba's international marketplace team has released and open-sourced CommerceAgentBench (GitHub repo Accio-org/CommerceAgentBench, already at 1.2k stars), a benchmark for testing whether AI Agents can actually get real e-commerce work done rather than just give correct answers. The benchmark contains 107 tasks covering the full chain—procurement, product listing, operations, fulfillment, and after-sales—distilled from real Alibaba international e-commerce data, with sources including 10 million SMB users and 1.6 million complete conversations. The best-performing agent currently completes only 61.68% of tasks.
Confirmed
- The benchmark went live and open-source on GitHub on August 25, repo Accio-org/CommerceAgentBench, 1.2k stars
- 107 real e-commerce tasks covering procurement, product listing, operations, fulfillment, and after-sales
- Task data distilled from real Alibaba international e-commerce data, involving 10 million SMB users and 1.6 million complete conversations
- Scoring is not based on the model's self-reports but on real action traces left in the environment: whether tags were correctly applied, drafts actually saved, calendar events created, and whether listings, orders, tracking numbers, and cross-system data are consistent
- Procurement tasks include pricing challenges: supplier quotes involve different currencies, Incoterms, surcharges, minimum order quantities, payment terms, and lead times; agents must normalize these into comparable landed costs, otherwise the cheapest supplier on paper may not be cheapest in reality
- The procurement test is set up as taking over a procurement inbox of 300 emails, mixing supplier quotes, quantity changes, delivery terms, payment info, withdrawn quotes, and similarly named suppliers, requiring agents to judge which information is valid and make procurement decisions rather than simply summarize emails
Why it matters
- Traditional benchmarks focus on whether a model gives correct answers; this one shifts to an Agent's actual execution ability, with scoring by operational outcomes better reflecting performance in real work scenarios
- The best agent completes only 61.68% of tasks, showing current Agents still have huge room for improvement in complex real e-commerce workflows, providing the industry with a new capability yardstick
2026-08-26 ~ 2026-08-26 · 5 related posts
Primary sources
- CommerceAgentBench Launch: Testing If AI Agents Can Complete Real E-commerce Tasks — alifcoder ·
- Best Agent Completes Only 61.68%: Alibaba's CommerceAgentBench Goes Live on GitHub — alifcoder ·
- CommerceAgentBench Grades Agents on Operational Traces, Not Self-Reports — and Hides BEC Fraud in the Inbox — alifcoder ·
- [source] CommerceAgentBench Launch: Testing If AI Agents Can Complete Real E-commerce Tasks — alifcoder · 2026-08-26
- Handed a 300-Email Procurement Inbox, the Agent Must Decide What's True and Buy — alifcoder · 2026-08-26
- Quotes in 6 Currencies, 5 Incoterms: Benchmark Forces Agents to Compute Real Landed Cost — alifcoder · 2026-08-26
- [source] CommerceAgentBench Grades Agents on Operational Traces, Not Self-Reports — and Hides BEC Fraud in the Inbox — alifcoder · 2026-08-26
- [source] Best Agent Completes Only 61.68%: Alibaba's CommerceAgentBench Goes Live on GitHub — alifcoder · 2026-08-26