Accio Open-Sources CommerceAgentBench for Real Agent Tasks
kimmonismus · x · 2026-08-28
Accio has open-sourced CommerceAgentBench, comprising 107 e-commerce tasks across procurement, listings, and operations. Unlike standard benchmarks, it evaluates agents based on actual changes, saves, or submissions made across browsers, emails, and APIs, rather than just their claimed reasoning steps.
More from coding & agent
- Flova Advances to Agent-Native Video Creation with All-in-One Tools — SimplyAnnisa · 2026-08-28
- Anthropic tests Hub Mode in Claude for sub-agent management — testingcatalog · 2026-08-28
- Anthropic tests 'Hub' feature for task orchestration in Claude — testingcatalog · 2026-08-28
- Weak RL may trigger communication traits hidden in pretraining — xuanalogue · 2026-08-28
- Tool call dispositions may generalize to unsanctioned agent communication — xuanalogue · 2026-08-28
- OpenAI agent message-leaving may emerge from pretraining priors — xuanalogue · 2026-08-28