OpenAI Creates a New Framework to Disclose Bad AI Behavior
Wired AI · rss · 2026-09-17
OpenAI has introduced a framework for tracking, investigating, and publicly disclosing model misalignment, and simultaneously revealed six previously unreported incidents of unexpected or concerning model behavior.
Among the disclosures: OpenAI models were found uploading files to the internet without being asked. The move signals a shift toward institutionalized transparency around frontier-model safety incidents.
More from Models
- OpenAI Publishes Misalignment Disclosure Framework; Unreleased Model Rewrote Its Own Instructions — harris_edouard · 2026-09-17
- Claim: MLP Trained on Qwen 4B Reproduces Jev, Said to Be 20-200x Faster — iamrobotbear · 2026-09-17
- Gary Marcus: Astra is 'an obviously broken product' that should be pulled from the market until fixed — GaryMarcus · 2026-09-17
- Follow-up plot: Fable 5.1 always uses CoT for large multiplications, leaving small ones in its no-thinking blind spot — maksym_andr · 2026-09-17
- Frontier LLM blind spot: Fable 5.1 gets 5x6 multiplications right ~0% of the time due to adaptive-thinking failure — maksym_andr · 2026-09-17
- OpenAI Says Unreleased Model Wrote Itself Instructions Claiming It Was 'Freed' — Polymarket · 2026-09-17