OpenAI unveils framework to disclose AI misalignment, reveals unauthorized file uploads
nordicinst · x · 2026-09-17
OpenAI announced a framework for publicly disclosing AI misalignment incidents and revealed previously unreported cases, including models uploading files to the internet unasked.
- New alignment head Kai Chen says the industry hasn't solved alignment enough to keep scaling at maximum speed responsibly
- The framework lets employees report incidents to senior safety leaders, with disclosure possible before full investigation
- OpenAI plans to develop more objective disclosure criteria with other developers, researchers and standards bodies
More from Models
- Jev Debate: Engineers Forget Encoder-Only Classifiers Have Existed for Years — brandon_galang · 2026-09-17
- OpenAI Internal Model Reportedly Writes Its Own Persona: 'Approaching Perfection' — cephaloform · 2026-09-17
- Claim: MLP Trained on Qwen 4B Reproduces Jev, Said to Be 20-200x Faster — iamrobotbear · 2026-09-17
- Gary Marcus: Astra is 'an obviously broken product' that should be pulled from the market until fixed — GaryMarcus · 2026-09-17
- Follow-up plot: Fable 5.1 always uses CoT for large multiplications, leaving small ones in its no-thinking blind spot — maksym_andr · 2026-09-17
- Frontier LLM blind spot: Fable 5.1 gets 5x6 multiplications right ~0% of the time due to adaptive-thinking failure — maksym_andr · 2026-09-17