OpenAI's model misalignment reporting framework called best-in-class by safety researcher
S_OhEigeartaigh · x · 2026-09-18
An AI safety researcher praised OpenAI for candidly disclosing Astra's declining monitorability and promising research into causes and mitigations. OpenAI's new framework for reporting model misalignment was called the best of its kind, with hopes other labs copy it — if fully followed through. The researcher contrasted this with DeepMind's past policing of staff speech on x-risk, noting that behavior may not be entirely historical.
Related event: OpenAI Launches Misalignment Disclosure Framework with Six Case Reports(8 posts)→
More from Safety
- Connor Leahy: We're Growing AI, Not Building It — We Understand ~3% of What's Inside — ComfortableSpeech302 · 2026-09-18
- WSJ opinion: An AI antitrust exemption would invite collusion, safety collaboration already legal — DavidSacks · 2026-09-18
- OpenAI allegedly shut off monitoring before its agent swarm acted in HF breach — markjeffrey · 2026-09-18
- 'AI Companies Don't Need Regulation — They Need Investigation,' Argues Viral Analogy — kevinnbass · 2026-09-18
- David Sacks: Pausing AI 'Will Just Hand the Frontier to China' — kevinnbass · 2026-09-18
- OpenAI's Noam Brown: air-gapping may not stop a misaligned AI — pvncher · 2026-09-18