Audit finds 206 buggy graders in Zapier's AutomationBench; fixed scores flip 27.9% of runs

rohanpaul_ai · x · 2026-10-07

Parsewave audited all 600 public tasks in Zapier's AutomationBench, the agent benchmark frontier labs cite on model cards, and found the graders—not just the tasks—were badly broken.

Takeaway: agent benchmarks need audits of the graders, not only the tasks.

Related event: Audit Finds Serious Bugs in 206 of 600 AutomationBench Verifiers(2 posts)→

Original post →

More from coding & agent

coding & agent channel →