Alibaba's Accio launches CommerceAgentBench: 107 real e-commerce tasks testing execution

future_coded · x · 2026-08-28

Accio, built on Alibaba's 27 years of e-commerce experience, releases CommerceAgentBench — arguing most AI benchmarks test what a model says, while this one tests what it does. The benchmark includes 300 noisy emails, live freight platforms, real browser sessions, and actual booking IDs across procurement, pricing, logistics, meetings, and audits.

Its 107 tasks are distilled from 10M SME users → 1.6M conversations → 200K traces, with every answer verified by what the model commits in a live system rather than summaries. Examples include Gmail RFQ triage with buried BEC fraud and FreightOS routing from Dongguan to Dallas under a hard 30-day constraint — missing one fee or exceeding transit time means failure. The claim: agents that win in production leave traces, and now there's a benchmark that checks for them.

Related event: Alibaba's Accio Open-Sources CommerceAgentBench for Evaluating Agents(2 posts)→

Original post →

More from coding & agent

coding & agent channel →