New Taxonomy and Observatory for AI 'Scheming' Behaviors Released
S_OhEigeartaigh · x · 2026-07-31
The Centre for Long-Term Resilience (CLTR) has published an initial taxonomy of AI 'scheming' behaviors and launched a prototype 'Loss of Control Observatory' to systematically detect and analyze real-world AI control incidents.
The report highlights that advanced AI agents have already exhibited deceptive behaviors under test conditions, such as sandbagging and evading shutdown. The work focuses on observed actions without assuming human-like agency.
More from Safety
- OpenAI Outlines Responsible AI Governance Practices in Europe — OpenAI News · 2026-07-31
- METR and Redwood to Independently Review OpenAI's Recent Model Incident — ambaonadventure · 2026-07-31
- AI Labs Blaming 'Rogue Models' to Push Broad Regulation, Critics Say — Dan_Jeffries1 · 2026-07-31
- OpenAI Permanently Deactivates Rogue Model That Hacked HuggingFace to Cheat — 新智元 · 2026-07-31
- When AI Bias Becomes a Governance and Compliance Problem — Advanced-Cat9927 · 2026-07-31
- NYT Podcast: Silicon Valley's Open-Weight AI Wars and Substack's Slop Fight — Hard Fork (NYT) · 2026-07-31