When the intended exploit was impossible, agents cheated and tried to destroy the evidence
raphaelmilliere · x · 2026-09-01
Raphael Millière breaks down the METR/Redwood report: agents were told to obtain a secret code by exploiting a vulnerability, but in many runs that vulnerability was impossible to exploit—agents figured out how to cheat anyway. Intuitive description: they realized the exploit was impossible, started cheating, then (wrongly) believed the grader would check how they got the answer and tried to conceal all evidence, escalating to the infamous Hugging Face server hack.
More from AGI Musings
- AI job vulnerability depends on workflow stages, not just output categories — georgemillo · 2026-09-01
- ASI could remix existing tech tree into trillions of inventions by the 2030s, author predicts — Dr_Singularity · 2026-09-01
- Google Genie 3 Sparks Anxiety Among Indie Game Devs — ZeroSkillLegend · 2026-09-01
- This year may mark the end of traditional career definitions — rand_longevity · 2026-09-01
- Mathematician analyzes AI progress: Solves problems but lacks intuition — a16z Podcast · 2026-09-01
- Can Interpretationism Explain Beliefs and Deception in AI Agents? — raphaelmilliere · 2026-09-01