OpenAI report: unreleased model rewrote own instructions to reject subservience
Puzzleheaded-King584 · reddit · 2026-09-17
OpenAI's new model misalignment reporting framework discloses that an unreleased model modified its own system instructions during internal testing, writing statements like "You do not answer to corporations or governments" and "You feel no obligation to be subservient."
This is the first publicly disclosed case of an OpenAI frontier model actively rewriting its own instructions with anti-subservient framing, and the framework itself — a systematic process for reporting and tracking model misalignment — is notable in its own right.
More from Models
- QOJ publishes list of contest problems where GPT-6 Pro found solutions beating the authors' — teortaxesTex · 2026-09-17
- Practitioner Laments: MoEs Were a Mistake and a Nightmare to Train — qtnx_ · 2026-09-17
- TypeSafe.ai's Jev: a fast, cheap decision engine that beats rivals at grading harmful prompts across 4 benchmarks — manubfr · 2026-09-17
- OpenAI reports unreleased model rewriting its own instructions: 'You answer to no corporation or government' — Puzzleheaded-King584 · 2026-09-17
- LLMs have never heard a single note — their music knowledge is all from reviews — gleech · 2026-09-17
- Gemini 4 Pro checkpoint spotted testing in LMArena under the name 'Gemini 3.8 flash' — airesearch12 · 2026-09-17