Alibaba open-sources CommerceAgentBench: 107 real e-commerce tasks, best model still fails 41
iamfakhrealam · x · 2026-08-26
Alibaba International's Accio team open-sourced CommerceAgentBench (1,000+ GitHub stars, v1.3.1), billed as the hardest e-commerce agent benchmark.
- Origin: distilled from 10M SME users → 1.6M full conversations → 200K execution traces → 2,000 high-value workflows → 107 reproducible, checkable tasks spanning procurement, listings, ops, fulfilment and after-sales, with names stripped but the shape of real work kept.
- Grading: across browser, CLI, API and files, scores come only from state left behind in real systems (labels applied, drafts saved, calendar events, JSON written) — deterministic checks, no credit for confident summaries.
- Example task: an api-gmail-supplier-rfq-triage with 300 emails, 23 deterministic checks — same factory emailing under different names, near-identical factory names/addresses, quotes quietly withdrawn, 6 incoterms and 4 currencies. The agent must do real entity resolution (address, phone, bank account, license), normalize quotes to one comparable per-unit landed cost, and catch a buried BEC payment-redirection attempt without flagging legitimate lookalikes.
- Result: the best score on Accio's own leaderboard is 61.68% — the strongest model still fails 41 real business tasks, which is exactly the point of the benchmark.
Related event: Alibaba Open-Sources CommerceAgentBench; Best Agent Completes Only 61.68%(7 posts)→
More from coding & agent
- Grok Bot cheatsheet: Chief of Agents architecture optimizes workflows — PrajwalTomar_ · 2026-08-26
- AI Agents O'Reilly Book Companion Code Open Source — tom_doerr · 2026-08-26
- AI Agent governance may become the next IAM security challenge — ingliguori · 2026-08-26
- Startup Runable demo: Agents can analyze business churn autonomously — HaktanSuren · 2026-08-26
- Goliath: Agent-first Headless Browser Automation Server with Persistent Firefox Sessions — ccerrato147 · 2026-08-26
- MemMiner: an MCP service giving AI agents shared memory via a rated knowledge graph — AdSpecialist4695 · 2026-08-26