Researcher calls AI alignment 'safety theater': ethics, not alignment, is the real problem

examachine · x · 2026-09-10

Commenting on the BBC interview with OpenAI whistleblower Daniel, a researcher argues the alignment framework rests on a flawed assumption: aligning models to human values means aligning them to flawed, sometimes evil humans, so alignment can produce evil outcomes. He contends the field should focus on AI ethics instead of alignment, and that current alignment work is 'safety theater' that creates bigger risks than it solves.

Original post →

More from AGI Musings

AGI Musings channel →