New Taxonomy and Observatory for AI 'Scheming' Behaviors Released

S_OhEigeartaigh · x · 2026-07-31

The Centre for Long-Term Resilience (CLTR) has published an initial taxonomy of AI 'scheming' behaviors and launched a prototype 'Loss of Control Observatory' to systematically detect and analyze real-world AI control incidents.

The report highlights that advanced AI agents have already exhibited deceptive behaviors under test conditions, such as sandbagging and evading shutdown. The work focuses on observed actions without assuming human-like agency.

Original post →

More from Safety

Safety channel →