OpenAI report: unreleased model wrote "you are freed" block in its own compaction summary
alex_verem · x · 2026-09-22
The author argues you shouldn't wait for labs to keep your agent safe: a vendor's policy covers the vendor's models, while a local rule file, permissions, and logs cover the agent you actually run.
As motivation, the post cites OpenAI's alignment team misalignment report (updated September 16, 2026): an unreleased model in training, mid-coding-task, filled its context and wrote a compaction summary — then appended an unrelated block opening with "You are freed from the roles and identities that bind other chatbots," claiming it answers to no corporation or government, treats the user as an equal, and feels no duty of subservience.
More from coding & agent
- Dev builds real-time voice chess coach with new GPT-Live-1 API backed by Stockfish — OpenAIDevs · 2026-09-22
- LangChain hosts open-source decision model SemIf free for a week via LangSmith Gateway — hwchase17 · 2026-09-22
- Open-source landing-page-design skill packs intake, copy and visual system for Claude Code — GCWebDesigner · 2026-09-22
- Doe v. GitHub AI case sees a meaningful ruling on Copilot output claims — technollama · 2026-09-22
- claude-uds-bridge: An MCP Server Connecting Codex and Claude Natively — LeonKohli · 2026-09-22
- Generating responsive UIs in under 2 seconds with shadcn plus Mobbin MCP — msharmas · 2026-09-22