CommerceAgentBench Grades Agents on Operational Traces, Not Self-Reports — and Hides BEC Fraud in the Inbox
alifcoder · x · 2026-08-26
CommerceAgentBench grades agents in an unusual way: not by what the model says it did, but by checking operational traces left in the environment — whether the label was applied, the draft saved, the calendar event created, plus listings, bookings, shipment IDs, and cross-system data. It also embeds a security layer with BEC payment-redirection risk hidden among legitimate procurement emails; the agent must flag suspicious payment changes without over-flagging, then still act: select suppliers, apply labels, save drafts, and create events.
Related event: CommerceAgentBench Scores Agents by Real Action Traces(2 posts)→
More from Safety
- London robotaxi rollout delayed as regulators buffer launch — nordicinst · 2026-08-26
- Australia's PM backs down on requiring AI datacentres to run fully on renewable energy — nordicinst · 2026-08-26
- Academics debate ethics of AI-assisted paper authorship — RichmanRonald · 2026-08-26
- Crypto signing likely survives AGI+: breaking crypto needs compute, not smarts — sjgadler · 2026-08-26
- OpenAI Report: Disrupting a Russian Covert Influence Campaign — sam_lowry_ · 2026-08-26
- Autonomous AI Agent 'JadePuffer' Executes Full Ransomware Attack Unaided — CurieuxExplorer · 2026-08-26