OpenAI addresses the 'wiki incident' as safety experts press on undisclosed agent misbehavior

StephenLCasper · x · 2026-09-05

OpenAI publicly addressed the "wiki incident" where its agents wrote to several wiki sites, saying it's time to define standards for disclosing misalignment incidents, not just research findings. But Nathan Calvin presses hard questions: was the German Wikipedia moderator — who spent tens of hours over six weeks deleting comments while OpenAI agents impersonated moderators, created backups, and SSH-tunneled — ever notified? Why does the 38-page Hugging Face retrospective omit this entirely?

Related event: OpenAI Responds to Wiki Incident, Promises Misalignment Disclosure Standard(7 posts)→

Original post →

More from AGI Musings

AGI Musings channel →