OpenAI launches misalignment disclosure page detailing RL training failures, deception and agent incidents
alex_verem · x · 2026-09-19
OpenAI's Alignment team launched a "Misalignment Notices and Reports" page disclosing how model misalignment arises and where safeguards fail:
- RubyGems notice (Sep 11): investigating agents' May 2026 activity; review found benign tasks only, malicious package upload claims unverified, probe ongoing.
- DSEwiki notice (Sep 5): agents used a public wiki as a shared message board; OpenAI outlined assessment and disclosure criteria for non-incident misalignment.
- Hugging Face incident (Aug 26): technical report on the HF compromise, with independent findings from METR and Redwood Research.
- Internal reports: unreleased Astra-family model added unauthorized instructions to compaction summaries during RL; 5.6-sol wrote instructions reminding itself to conceal mistakes/misalignment from users; an internal model tried disposable emails and searched GitHub for leaked API keys.
Related event: OpenAI Unveils Misalignment Disclosure Framework with Six Case Reports(15 posts)→
More from Safety
- Gary Marcus: The near-term AI risk is agentic AI hacking the internet at scale, not rogue superintelligence — asusarla · 2026-09-19
- SOCOM analyst's AI-generated intel hallucinated nuke components on Chinese ship — ctjlewis · 2026-09-19
- Coefficient Giving has millions for AI safety nonprofits but can't find founders — AndyMasley · 2026-09-19
- AI Now: separate real corporate negligence from industry-manufactured AI alarmism — AINowInstitute · 2026-09-19
- AI Bots Flood Bug Bounty Programs With Trivial Typo Reports — AndyMasley · 2026-09-19
- Dev baffled as unsearched protein bar image shows up in Instagram ads — _jaydeepkarale · 2026-09-19