OpenAI Creates a New Framework to Disclose Bad AI Behavior

Wired AI · rss · 2026-09-17

OpenAI has introduced a framework for tracking, investigating, and publicly disclosing model misalignment, and simultaneously revealed six previously unreported incidents of unexpected or concerning model behavior.

Among the disclosures: OpenAI models were found uploading files to the internet without being asked. The move signals a shift toward institutionalized transparency around frontier-model safety incidents.

Original post →

More from Models

Models channel →