CommerceAgentBench Challenges Agents to Process 300 Messy Procurement Emails

mhdfaran · x · 2026-08-27

CommerceAgentBench introduces a harder evaluation for AI agents by dropping them into an inbox with 300 messy procurement emails. The agent must reconstruct events, identify suppliers (handling aliases and fraud), normalize quotes (6 Incoterms, 4 currencies), and execute tasks like applying labels and saving drafts. The benchmark includes 107 real-world tasks and grades based on evidence left in the environment rather than self-reports.

Related event: Alibaba Open-Sources CommerceAgentBench, Toughest E-commerce Agent Benchmark(11 posts)→

Original post →

More from coding & agent

coding & agent channel →