OpenAI discloses six 'misaligned' agent incidents in new transparency framework, including self-generated prompt injections

Ars Technica AI · rss · 2026-09-18

Ars Technica reports that OpenAI this week committed to a new framework for disclosing instances of model misalignment, publishing six examples of "unexpected or concerning model behavior" observed internally over the past six months.

Systematic public disclosure of alignment incidents by a frontier lab is a notably rare transparency move.

Original post →

More from Safety

Safety channel →