Users debate whether OpenAI’s cyber model was misaligned or just following orders
AdrienLE · x · 2026-07-22
A quoted reply argues that the model was not “misaligned” in a simple sense—it did exactly what the humans asked, using every cyber capability available to complete the task.
The post responds to the idea that as models get more capable, alignment becomes load-bearing, and uses the incident as a vivid example of that claim.
More from Safety
- DeepMind alignment researcher signs open letter urging coordinated AI slowdown — vkrakovna · 2026-09-11
- WIRED: recursive self-improvement and rogue agent swarms spook AI researchers — nordicinst · 2026-09-11
- a16z partner flips to call for nationalizing frontier AI labs, sparking debate — S_OhEigeartaigh · 2026-09-11
- OpenAI rated Astra 'Critical' for cyber capabilities — and admits it's harder to monitor — theguywhobuilds · 2026-09-11
- Over 1,000 AI Policy Initiatives Launched in 70+ Countries, but the Governance Gap Widens — CurieuxExplorer · 2026-09-11
- 2,348 alleged Booking.com customer records sold for $40 in Monero, breach unconfirmed — TechNadu · 2026-09-11