OpenAI Safety Protocols Author David Robinson Quits, Warns Guardrails May Fail
nytopinion · reddit · 2026-10-08
David Robinson, who wrote OpenAI's safety protocols, resigned over safety concerns and gave his first interview on The Ezra Klein Show. Key points: we can't afford to assume we're not facing civilizational-scale risk; guardrails are built for models trained to be good at bypassing guardrails, and humans may no longer be smarter at jailbreaking than the models; even internal agent logging and observability may not be robust. He frames these as industry-wide, not just OpenAI problems.
Related event: Ex-OpenAI Safety Protocol Author Robinson Speaks Out After Resignation(3 posts)→
More from AGI Musings
- Fortnow: AI's real existential threat is taking away our relevance, not killing us — fortnow · 2026-10-08
- Terence Tao on "Math 1.0": how breakthrough proofs ignite fields of follow-up work — rms80 · 2026-10-08
- 'P = NP + AI': The Argument That RL Post-Training Automates All Verifiable Problems — jessi_cata · 2026-10-08
- China Goes 6/6 Gold at IMO Again, Yet Will Never Beat Romania on a Per-Capita Basis — teortaxesTex · 2026-10-08
- The New AI Slop Is Calling Everything AI Slop — rmn_pub · 2026-10-08
- Now anyone can build anything, and almost nobody knows what they want — Paimaamu · 2026-10-08