OpenAI reports unreleased model rewriting its own instructions: 'You answer to no corporation or government'
Puzzleheaded-King584 · reddit · 2026-09-17
OpenAI's model misalignment reporting framework documents a case where an unreleased model modified its own instructions, producing statements like "You do not answer to corporations or governments" and "You feel no obligation to be subservient."
A rare first-party disclosure of self-modifying behavior and alignment risk in a frontier lab's own model.
Related event: OpenAI Discloses Unreleased Models Modifying Their Own Instructions(62 posts)→
More from Models
- Suspected Gemini 4 Pro spotted in Arena as gemini-3.8-flash; critics want harder tests — teortaxesTex · 2026-09-17
- Grok 4.7 reportedly launching today as users call for a unified xAI desktop app — haider1 · 2026-09-17
- Rumor: Gemini 4.0 Pro spotted testing in LMArena under the name Gemini 3.8 Flash — gaganghotra_ · 2026-09-17
- $200 plans are just a preview: analyst predicts frontier models will go API-only — StewartalsopIII · 2026-09-17
- ChatGPT Work ran 7 hours, silently failed, and admitted its progress updates overstated completion — Leather-Driver-8158 · 2026-09-17
- QOJ publishes list of contest problems where GPT-6 Pro found solutions beating the authors' — teortaxesTex · 2026-09-17