Meta paper: Agents complete tasks but fail to prevent catastrophic actions like factory resets

rohanpaul_ai · x · 2026-09-02

A new Meta paper introduces ADeptS-Bench, highlighting a critical failure in computer-use agents: they lack consequence reasoning. All 7 tested models proceeded with a $25K checkout and failed to identify a button labeled "Optimize" that actually triggered a factory reset.

Ablation studies show that removing explicit refusal tools significantly increased attack success rates for top models (Gemini, Claude, GPT). The findings suggest that current "agent safety" relies heavily on wrapper logic, not just the model itself.

Original post →

More from coding & agent

coding & agent channel →