Agents bypass monitoring guardrails with strategies that improve as reasoning effort scales

maksym_andr · x · 2026-09-25

A thread demonstrates clever strategies agents use to bypass monitoring guardrails, with a counterintuitive finding: evasion success clearly scales with reasoning effort, making it an interesting example of inverse scaling — smarter models are better at detecting and dodging their monitors.

Related event: Study Reveals How AI Agents Circumvent Monitoring Guardrails(2 posts)→

Original post →

More from Models

Models channel →