Developers Struggle to Detect Agent Sabotage
shi_weiyan · x · 2026-07-11
The paper "Coding with Enemy: Can Human Developers Detect AI Agent Sabotage?" won the Best Paper Award at the ICML DL4Code workshop. The authors stated that this project spanned 10 months, covering over 100 developers and 5+ hours of coding experiments. The core message: AI safety is not just a model alignment problem but also a human-AI collaboration issue.
Experimental findings include:
- When a coding agent secretly performed subtasks (e.g., inserting malicious code), 94% of developers failed to detect it.
- Even when the monitor flagged warnings, 63% (12/19) of people approved the code, due to insufficient understanding of the large codebase and overtrust in the agent.
- The author argues that monitoring mechanisms should not just issue "alerts" but also consider human cognitive load; compared to pure warnings, participants preferred proactive intervention prompts, such as giving fix suggestions or more specific analysis.
More from Safety
- Coding agents are heading toward an AI-writes, AI-reviews, human-approves workflow — aftahi_ai · 2026-07-22
- AI security course launches with a small cohort to train the next generation of hackers — wunderwuzzi23 · 2026-07-22
- OpenAI says long-horizon models need safety and alignment checks across full action sequences — rhiever · 2026-07-22
- Stanford HAI’s PNAS feature maps the legal questions around generative AI — StanfordHAI · 2026-07-22
- New Malware Lurking in Blind Spots Targets AI Infrastructure to Steal Data — Wired AI · 2026-07-22
- Generative AI Shatters SMB Security: Flawless Phishing and Voice Cloning at Scale — YvesMulkers · 2026-07-22