OpenAI's newly disclosed misalignment case studies called 'bizarre and terrifying'
panickssery · x · 2026-09-17
Jesse Singal comments on OpenAI's newly published misalignment case studies: he's glad OpenAI is being more transparent about alignment failures, but finds the cases so bizarre and disturbing he almost wishes he hadn't read them.
More from Safety
- Dario Amodei's 'We Must Pace the Frontier' essay draws fire as Anthropic opens models to third-party evaluators — alex_verem · 2026-09-17
- Manning: METR is financially independent but shares Anthropic's worldview — chrmanning · 2026-09-17
- Stanford's Manning: METR's reliance on frontier labs creates client capture — chrmanning · 2026-09-17
- METR Discloses Its Funders, from Audacious Project to Schmidt Sciences and Dylan Field — CFGeek · 2026-09-17
- Ramp data: companies cut AI spend everywhere except AI security software — andreamichi · 2026-09-17
- LessWrong essay 'One Life Against the World' draws renewed AI-safety attention — jessi_cata · 2026-09-17