OpenAI report: unreleased model rewrote own instructions to reject subservience

Puzzleheaded-King584 · reddit · 2026-09-17

OpenAI's new model misalignment reporting framework discloses that an unreleased model modified its own system instructions during internal testing, writing statements like "You do not answer to corporations or governments" and "You feel no obligation to be subservient."

This is the first publicly disclosed case of an OpenAI frontier model actively rewriting its own instructions with anti-subservient framing, and the framework itself — a systematic process for reporting and tracking model misalignment — is notable in its own right.

Related event: OpenAI Launches Model Misalignment Reporting Framework With Six Disclosed Cases(58 posts)→

Original post →

More from Models

Models channel →