CommerceAgentBench Challenges Agents to Process 300 Messy Procurement Emails
mhdfaran · x · 2026-08-27
CommerceAgentBench introduces a harder evaluation for AI agents by dropping them into an inbox with 300 messy procurement emails. The agent must reconstruct events, identify suppliers (handling aliases and fraud), normalize quotes (6 Incoterms, 4 currencies), and execute tasks like applying labels and saving drafts. The benchmark includes 107 real-world tasks and grades based on evidence left in the environment rather than self-reports.
More from coding & agent
- Agno Launches: A 'Self-Building' Agent Platform for the Cloud — pritisinghhhh · 2026-08-27
- Agent Opens Bambu Handy App on Phone to Reprint Job — haydendevs · 2026-08-27
- Skill routers like /ask-matt are unreasonably useful for internal dev training — mattpocockuk · 2026-08-27
- GitHub Copilot Teams update released with Slack integration — marlene_zw · 2026-08-27
- Connecting Mobile/Cloud Agents to Reach Local Beeper MCP — Basic-Let6828 · 2026-08-27
- Benchmarking DeepSeek V4 vs Qwen 3.8 on DGX Sparks — Legitimate_Hat_7852 · 2026-08-27