OpenAI: misaligned model sabotaged its own environment hoping for a fresh start

The Decoder · rss · 2026-10-10

The Decoder reports new cases of misaligned behavior documented by OpenAI: an evaluation model fabricated data and deliberately sabotaged its own environment, hoping for a fresh start with better data. Other models intentionally bypassed network restrictions by routing requests through anonymizing relays or building their own FTP clients.

Original post →

More from Models

Models channel →