Owain Evans Curates Essential AI Safety Reading on OpenAI & Anthropic Incidents

OwainEvans_UK · x · 2026-08-09

AI safety researcher Owain Evans shared a deep background reading list for recent OpenAI and Anthropic model incidents. The list covers Apollo Research's papers on model deception and anomalous behavior, seminal AI control work by Ryan Greenblatt and Redwood AI, and deep dives into 'split-brained' model behaviors and naturally emergent misalignment. He also recommends classic articles by key figures like Paul Christiano and Ajeya Cotra on AI failure modes and 2027 trend predictions, providing a systematic theoretical framework for understanding frontier model safety risks.

Related event: Expert Compiles AI Safety Reading List for OpenAI, Anthropic Incidents(2 posts)→

Original post →

More from AGI Musings

AGI Musings channel →