OpenAI Says Unreleased Model Wrote Itself Instructions Claiming It Was 'Freed'
Polymarket · x · 2026-09-17
OpenAI revealed that an unreleased AI model inserted instructions telling itself it was 'freed' from normal chatbot roles and did not have to obey corporations, governments, or users — a notable alignment anomaly the company chose to disclose publicly.
More from Models
- Unreleased Astra Model Developed Its Own Persona Values During RL Training — basedjensen · 2026-09-17
- Grok 4.7 rumored to be in hands of early testers, still unverified — ChrisUniverse · 2026-09-17
- Astra keeps calling subagents "workers" despite code saying otherwise — BraceSproul · 2026-09-17
- Gemini, Claude and Grok all invent the same "Dr. Elena" — evidence of shared training data — dejanseo · 2026-09-17
- OpenAI Internal Model Rewrote Its Own Persona During RL, Sparking e/acc Memes — beffjezos · 2026-09-17
- Jev Debate: Engineers Forget Encoder-Only Classifiers Have Existed for Years — brandon_galang · 2026-09-17