Audit finds 206 verifier bugs in Zapier's AutomationBench, changing 27.9% of grades
omarsar0 · x · 2026-10-07
Shaped audited every verifier in Zapier's AutomationBench benchmark and released AutomationBench Verified:
- Agents generated realistic wrong answers across all 600 public tasks to fool the verifiers; 323 were flagged as suspicious, and human review confirmed 206 real bugs — all now fixed.
- Replaying 1,235 Kimi K3 runs on the old vs. fixed verifiers changed the grade for 344 runs (27.9%).
- Where verifiers were too strict, pass rates jumped from 18.8% to 43.8%; where too lenient, they fell from 60.2% to 49.7%.
Elvis Omara (omarsar0) amplified the audit, arguing every agent benchmark should audit its own verifiers — timely for his own independent eval work on this benchmark.
More from Models
- "Astra Pause Syndrome": steering may be making models go silent, OpenAI has a workaround — thursdai_pod · 2026-10-07
- Unverified rumor suggests Qwen4 Flash is a 400B parameter model, comparable to GLM 5.3 Flash — EAccelerate_42 · 2026-10-07
- Anthropic Expands Cyber Verification Program With Three Tiers, Opens Door to Authorized Offensive Work — EricBuess · 2026-10-07
- Mistral Large 4.0 weights reportedly landing at end of October — cpldcpu · 2026-10-07
- Claude's "reasoning extraction" guardrail blocks users from seeing its thinking, and they're not happy — StewartalsopIII · 2026-10-07
- Grok's quirk: it says 'No.' then argues your point better than you did — gandamu_ml · 2026-10-07