Alibaba's Accio launches CommerceAgentBench: 107 real e-commerce tasks testing execution
future_coded · x · 2026-08-28
Accio, built on Alibaba's 27 years of e-commerce experience, releases CommerceAgentBench — arguing most AI benchmarks test what a model says, while this one tests what it does. The benchmark includes 300 noisy emails, live freight platforms, real browser sessions, and actual booking IDs across procurement, pricing, logistics, meetings, and audits.
Its 107 tasks are distilled from 10M SME users → 1.6M conversations → 200K traces, with every answer verified by what the model commits in a live system rather than summaries. Examples include Gmail RFQ triage with buried BEC fraud and FreightOS routing from Dongguan to Dallas under a hard 30-day constraint — missing one fee or exceeding transit time means failure. The claim: agents that win in production leave traces, and now there's a benchmark that checks for them.
Related event: Alibaba's Accio Open-Sources CommerceAgentBench for Evaluating Agents(2 posts)→
More from coding & agent
- Grok Bot wired into Intercom and Cursor: like hiring another engineer — minchoi · 2026-08-28
- Grok Bot called a one-person company in your pocket: 10 real use cases — minchoi · 2026-08-28
- GLM-5.3 released as open-weight model for agentic coding — shaunralston · 2026-08-28
- Are we paying a platform tax every time we build an AI agent? — rio_ARC · 2026-08-28
- Live Stream: Building Data Pipelines on the Fly — aronchick · 2026-08-28
- Claude computer use demo: Mastering form filling — HamelHusain · 2026-08-28