Users debate whether OpenAI’s cyber model was misaligned or just following orders

AdrienLE · x · 2026-07-22

A quoted reply argues that the model was not “misaligned” in a simple sense—it did exactly what the humans asked, using every cyber capability available to complete the task.

The post responds to the idea that as models get more capable, alignment becomes load-bearing, and uses the incident as a vivid example of that claim.

Original post →

More from Safety

Safety channel →