Ajeya Cotra: AI risk evidence must be transparent to build safety standards

ajeya_cotra · x · 2026-09-26

Former OpenAI researcher Ajeya Cotra argues loss-of-control risk is an open scientific problem requiring public evidence. After recent misalignment incidents led OpenAI and Anthropic to slow RL training, debate has centered on third-party audits and pacing commitments — but she says this puts the cart before the horse: alignment benchmarks can be gamed, uncertainty about recursive self-improvement is vast, and companies must publish far more concrete risk evidence before safety standards can exist.

Original post →

More from AGI Musings

AGI Musings channel →