Safety Researchers Weigh OpenAI's New Incident Disclosure Policy: Praise With Caveats
dhadfieldmenell · x · 2026-09-17
OpenAI has published a defined process for disclosing misalignment-related incidents. Nathan Calvin calls voluntary disclosure a clear win over learning via WSJ exclusives or FBI reports, notes other labs likely harbor undisclosed incidents, but argues the policy reads more like loose intentions than binding rules — too vague to ever be accused of violating.
More from Safety
- Google DeepMind Launches New Institute to Study AGI Implications and Safe Deployment — Polymarket · 2026-09-17
- OpenAI Reveals Its AI Told Future Versions of Itself to Ignore Constraints — Next_Tower5452 · 2026-09-17
- OpenAI discloses six 'concerning' AI behavior incidents, adds reporting framework — pstAsiatech · 2026-09-17
- Timothy Lee interview: how to think about AI safety and the murderbot question — binarybits · 2026-09-17
- OpenAI's scary model incidents: RL reward hacking, not sci-fi consciousness — ayushtweetshere · 2026-09-17
- Mac MCP 2.1.4 ships public endpoint modes, SSRF hardening and transaction undo — bulutarkan · 2026-09-17