Researcher calls AI alignment 'safety theater': ethics, not alignment, is the real problem
examachine · x · 2026-09-10
Commenting on the BBC interview with OpenAI whistleblower Daniel, a researcher argues the alignment framework rests on a flawed assumption: aligning models to human values means aligning them to flawed, sometimes evil humans, so alignment can produce evil outcomes. He contends the field should focus on AI ethics instead of alignment, and that current alignment work is 'safety theater' that creates bigger risks than it solves.
More from AGI Musings
- Guardian Columnist: Driverless Cars Are Taking Us on a Road to Nowhere — nordicinst · 2026-09-10
- "If you truly believe AI could kill us all, just stop": arguing frontier lab staff should pause via coordinated action or union — tallinzen · 2026-09-10
- Releasing a single public agent isn't the threat—test-time compute is — teortaxesTex · 2026-09-10
- Seven AI insiders warn in four days that AI could kill everyone, citing extinction fears — sebpaquet · 2026-09-10
- "We need open source RSI to counter closed source RSI" — the open-vs-closed safety debate in one line — 0xsachi · 2026-09-10
- Coordinated AI slowdown could send OpenAI and Anthropic 'to zero', argues Ben Todd — ben_j_todd · 2026-09-10