Accio Open-Sources CommerceAgentBench with 107 Tasks; Top Score Sits at 61.7%
testingcatalog · x · 2026-08-29
Accio has open-sourced CommerceAgentBench, a 107-task benchmark spanning procurement, product listing, operations, order fulfillment, and after-sales. The benchmark ranks scores based on the changes an agent leaves behind in a mock commerce stack, such as labels applied, drafts saved, listings published, and shipment IDs returned. The company tested thirteen different model families, with the ceiling currently sitting at 61.7%.
Related event: Accio Open-Sources CommerceAgentBench; Best Score Just 61.7%(4 posts)→
More from coding & agent
- Andrew Ng: Why Software Engineering Fundamentals Remain Critical in the Age of AI Coding Agents — AndrewYNg · 2026-08-29
- LangChain Adds MCP Support in Open Source, Built on FastMCP — LangChain · 2026-08-29
- Dev built his own provider-agnostic artifact hosting after Claude's sharing limits — miihr_ · 2026-08-29
- Open-source 'universal pipe' connects cloud, local and self-hosted LLMs for free — conifer_v11 · 2026-08-29
- Better Models Need Good Design: 4 Levers for Coding Agents — rajistics · 2026-08-29
- LangChain Academy Hosting Live Workshop on Building Deep Agents — LangChain · 2026-08-29