OpenAI's model misalignment reporting framework called best-in-class by safety researcher

S_OhEigeartaigh · x · 2026-09-18

An AI safety researcher praised OpenAI for candidly disclosing Astra's declining monitorability and promising research into causes and mitigations. OpenAI's new framework for reporting model misalignment was called the best of its kind, with hopes other labs copy it — if fully followed through. The researcher contrasted this with DeepMind's past policing of staff speech on x-risk, noting that behavior may not be entirely historical.

Related event: OpenAI Launches Misalignment Disclosure Framework with Six Case Reports(8 posts)→

Original post →

More from Safety

Safety channel →