CommerceAgentBench hits 1K stars: 107 e-commerce agent tasks distilled from 1.6M real conversations
VibeMarketer_ · x · 2026-09-02
CommerceAgentBench is a benchmark for commerce agents that has passed 1,000 GitHub stars and is becoming the default answer to whether an agent is ready for e-commerce work.
- Its 107 tasks span procurement, product listings, operations, fulfillment, and after-sales, mostly forcing cross-system execution: reading a messy inbox or quotes, then finishing the job in a browser, calendar, or supplier platform.
- Tasks come from real usage: Accio runs an AI agent for 10 million SMEs, and these 107 were cut from 1.6 million real conversations and 200K execution traces, backed by Alibaba's 27 years in e-commerce.
- The benchmark assumes humans keep making decisions, testing only whether an agent can handle execution underneath them.
More from coding & agent
- LunagraphHQ Launches: Visual Editor for React — floguo · 2026-09-02
- Rippling Launches AI Agent to Automatically Resolve Tedious IT Issues — andreisavu · 2026-09-02
- Fable 5.1 released: Agentic coding tasks solved at half the cost — airesearch12 · 2026-09-02
- Nested traces are outdated; observability tools should focus on chat views and failure surfacing — _ScottCondron · 2026-09-02
- Lovable on Model Routing: Moving Beyond the Model Picker — tylerbruno05 · 2026-09-02
- Backend Consensus Shift: Use Effect Instead of Plain TypeScript — mattpocockuk · 2026-09-02