Parsewave audit finds 206 serious verifier flaws in 600 public AutomationBench tasks
rohanpaul_ai · x · 2026-10-07
Parsewave audited AutomationBench — the Zapier-built agent benchmark featured on model cards like Gemini 4 Argon, Claude Opus 5.5, GPT-6 Astra, and GPT-6.1 Sol. On the 600 public tasks, agent-assisted adversarial submissions flagged 323 tasks; manual review confirmed 206 serious verifier issues (117 rejected as false positives), all now fixed, with detailed reasoning published for 100 findings. Experiments show the fixes materially change scores for an open-weights frontier model — a caution for reading agent benchmark leaderboards.
More from Research
- NeuroAI manifesto: applying AI scaling laws to BCIs, the bitter lesson for the brain — w1kke · 2026-10-07
- LeanLean benchmark: Opus 5.5 scores 64.3% compressing Lean proofs, GPT 6.1 Sol only 39.9% — ChrSzegedy · 2026-10-07
- PersistBench (NeurIPS Spotlight): 4D foundation models can see but not remember — weichiuma · 2026-10-07
- COLM 2026 poster: Differentiable Faithfulness Alignment for Cross-Model Circuit Transfer — boknilev · 2026-10-07
- AI-Read Gold Electrodes Detect Molecular Chirality One Molecule at a Time — Brighter-Side-News · 2026-10-07
- BinkBench: A No-Cap Agent Benchmark for Video Quality and Compression, Seeking Testers — -MaskNinja- · 2026-10-07