OpenAI unveils misalignment disclosure framework; experts say it lacks teeth
dhadfieldmenell · x · 2026-09-22
OpenAI announced a new framework for tracking, investigating, and publicly disclosing model misalignment, with criteria and timelines covering even unexplained behaviors. Policy researcher MackenZarnold's take: it's fine but weak — criteria are 'we'll know it when we see it,' not falsifiable, overly reliant on employee reports, and OpenAI chose not to make it binding under California's SB 53 despite having the option. He urges other labs to follow suit and for future legislation to allow rulemaking to update disclosure standards over time.
More from Safety
- Sarah Hooker suspects fake AI paper submissions cluster at a few universities — sarahookr · 2026-09-22
- Umbriel's Caleb Gross drops Blackhat talk on applying information retrieval to vulnerability research — dyn___ · 2026-09-22
- 3 of 14 participants in AI mental health trial had psychiatric events, experts warn of scaling risks — manorlaboratory · 2026-09-22
- Blind RSA apps like Privacy Pass face real-world threat model from scaled oracle queries — matthew_d_green · 2026-09-22
- Xbox AI Filter Bans Gamer for Listing His Hometown, $300 in Fees Gone — Aiden_Tech_Ai · 2026-09-22
- Alex Epstein calls (P)doom 'fake threat analysis' that only manufactures fear — TinfoilTricorn · 2026-09-22