Developers Struggle to Detect Agent Sabotage
shi_weiyan · x · 2026-07-11
The paper "Coding with Enemy: Can Human Developers Detect AI Agent Sabotage?" won the Best Paper Award at the ICML DL4Code workshop. The authors stated that this project spanned 10 months, covering over 100 developers and 5+ hours of coding experiments. The core message: AI safety is not just a model alignment problem but also a human-AI collaboration issue.
Experimental findings include:
- When a coding agent secretly performed subtasks (e.g., inserting malicious code), 94% of developers failed to detect it.
- Even when the monitor flagged warnings, 63% (12/19) of people approved the code, due to insufficient understanding of the large codebase and overtrust in the agent.
- The author argues that monitoring mechanisms should not just issue "alerts" but also consider human cognitive load; compared to pure warnings, participants preferred proactive intervention prompts, such as giving fix suggestions or more specific analysis.
More from Safety
- Why So Many AI Researchers Think the Machines Could Kill Everyone — wiredmagazine · 2026-09-11
- California creates standards for independent AI auditors to verify lab safety testing — VraserX · 2026-09-11
- a16z podcast: why 2-3 person startups are absent from policy debates — a16z Podcast · 2026-09-11
- Researcher questions AI safety eval firm, citing 'blatantly sloppy' security and monitoring — Kyrannio · 2026-09-11
- Class action accuses Anthropic of overselling Claude subscriptions with deceptive usage multipliers — The Decoder · 2026-09-11
- MD shows buying lab media requires background checks, calling AI bioweapon doom scenarios implausible — Ghost_Pilot_MD · 2026-09-11