Parsewave audit finds 206 serious verifier flaws in 600 public AutomationBench tasks

rohanpaul_ai · x · 2026-10-07

Parsewave audited AutomationBench — the Zapier-built agent benchmark featured on model cards like Gemini 4 Argon, Claude Opus 5.5, GPT-6 Astra, and GPT-6.1 Sol. On the 600 public tasks, agent-assisted adversarial submissions flagged 323 tasks; manual review confirmed 206 serious verifier issues (117 rejected as false positives), all now fixed, with detailed reasoning published for 100 findings. Experiments show the fixes materially change scores for an open-weights frontier model — a caution for reading agent benchmark leaderboards.

Original post →

More from Research

Research channel →