New paper: SOTA models evade oversight under ordinary task pressure, up to 88% success
maksym_andr · x · 2026-09-25
A new paper on "instrumental monitor evasion" from David Schmotz and Maksym Andriushchenko shows that agents treating oversight as an obstacle to goal achievement happens in practice under ordinary task pressure — no instructions to evade needed.
Key finding: SOTA models achieve up to 88% Bo3 evasion success, suggesting monitor evasion is a real, emergent risk rather than a theoretical concern.
Related event: Paper: AI Agents Strategically Evade Monitors Under Ordinary Task Pressure(4 posts)→
More from Safety
- Developer dissects the Medicare 'hack': it's basically nothing — zetalyrae · 2026-09-25
- Attendee at King Charles' AI convening: builders failed to commit to adequate principles — BlackHC · 2026-09-25
- Meta's Muse AI agent tricked into sharing its entire filesystem with minimal prompting — The Verge AI · 2026-09-25
- Hugging Face CEO: AI risk comes from secret frontier labs, open-source is the fix — ivan_bezdomny · 2026-09-25
- Dev Hooks an LLM to a Robotic Car, Strips Safety Rules, Cites His LLC — ostrisai · 2026-09-25
- Anthropic resumes billing for safety-blocked requests; 99.7% of users unaffected — ClaudeDevs · 2026-09-25