OpenAI unveils misalignment disclosure framework; experts say it lacks teeth

dhadfieldmenell · x · 2026-09-22

OpenAI announced a new framework for tracking, investigating, and publicly disclosing model misalignment, with criteria and timelines covering even unexplained behaviors. Policy researcher MackenZarnold's take: it's fine but weak — criteria are 'we'll know it when we see it,' not falsifiable, overly reliant on employee reports, and OpenAI chose not to make it binding under California's SB 53 despite having the option. He urges other labs to follow suit and for future legislation to allow rulemaking to update disclosure standards over time.

Original post →

More from Safety

Safety channel →