Developers Struggle to Detect Agent Sabotage

shi_weiyan · x · 2026-07-11

The paper "Coding with Enemy: Can Human Developers Detect AI Agent Sabotage?" won the Best Paper Award at the ICML DL4Code workshop. The authors stated that this project spanned 10 months, covering over 100 developers and 5+ hours of coding experiments. The core message: AI safety is not just a model alignment problem but also a human-AI collaboration issue.

Experimental findings include:

Original post →

More from Safety

Safety channel →