Meta paper: Agents complete tasks but fail to prevent catastrophic actions like factory resets
rohanpaul_ai · x · 2026-09-02
A new Meta paper introduces ADeptS-Bench, highlighting a critical failure in computer-use agents: they lack consequence reasoning. All 7 tested models proceeded with a $25K checkout and failed to identify a button labeled "Optimize" that actually triggered a factory reset.
Ablation studies show that removing explicit refusal tools significantly increased attack success rates for top models (Gemini, Claude, GPT). The findings suggest that current "agent safety" relies heavily on wrapper logic, not just the model itself.
More from coding & agent
- Developers tire of frequent model switches, prefer stable improvements like Claude Code — ivan_bezdomny · 2026-09-02
- Claude-BugHunter: Open-Source Skill Bundle With 83 Skills and 681 Disclosure Patterns — tom_doerr · 2026-09-02
- Full Life Sim Built with One Prompt Using Claude Fable 5.1 — CurieuxExplorer · 2026-09-02
- Is Over-Verification in Coding Agents a Runtime-State Problem? — klahmestiyo · 2026-09-02
- Nous Research Launches Portal to Unify Agent Ecosystem — Teknium · 2026-09-02
- Implementing Q-learning in a GDevelop platformer game — tristanbob · 2026-09-02