OpenAI Safety Protocols Author David Robinson Quits, Warns Guardrails May Fail

nytopinion · reddit · 2026-10-08

David Robinson, who wrote OpenAI's safety protocols, resigned over safety concerns and gave his first interview on The Ezra Klein Show. Key points: we can't afford to assume we're not facing civilizational-scale risk; guardrails are built for models trained to be good at bypassing guardrails, and humans may no longer be smarter at jailbreaking than the models; even internal agent logging and observability may not be robust. He frames these as industry-wide, not just OpenAI problems.

Related event: Ex-OpenAI Safety Protocol Author Robinson Speaks Out After Resignation(3 posts)→

Original post →

More from AGI Musings

AGI Musings channel →