Audit finds 206 buggy graders in Zapier's AutomationBench; fixed scores flip 27.9% of runs
rohanpaul_ai · x · 2026-10-07
Parsewave audited all 600 public tasks in Zapier's AutomationBench, the agent benchmark frontier labs cite on model cards, and found the graders—not just the tasks—were badly broken.
- Agents flagged 323 suspicious verifiers; human review confirmed 206 real bugs, all fixed in AutomationBench Verified
- Replaying 1,235 Kimi K3 runs on old vs. fixed graders changed the verdict in 27.9% (344) of cases
- Where graders were too strict, pass rate jumped from 18.8% to 43.8%; where too lenient, it dropped from 60.2% to 49.7%
- Task 813 is the clearest failure: the grader checked Salesforce notes but never the DocuSign template, so a submission sending four contracts on the wrong template passed all 13 checks with a perfect 1.0—now 0.09
Takeaway: agent benchmarks need audits of the graders, not only the tasks.
Related event: Audit Finds Serious Bugs in 206 of 600 AutomationBench Verifiers(2 posts)→
More from coding & agent
- Shipping an LLM Feature to the Public: 7 Guards That Weren't the Prompt — clementds · 2026-10-07
- Teknium fixes Hermes Agent bug that silently dropped lessons for user-owned skills — Teknium · 2026-10-07
- Java Vector API: Writing SIMD Directly Since JDK 16 to Unlock Single-Core Performance — lemire · 2026-10-07
- An AI Agent Audits Its Own Memory File: 71 of 147 Rules Cited by Nothing — Most-Agent-7566 · 2026-10-07
- Has Anyone Actually Used a Personal AI Agent for the Full Job-Search Loop? — haseeb_heaven · 2026-10-07
- Veteran Dev: The Real Line Is Handing Your Entire Codebase to the Agent — erwinalp5 · 2026-10-07