Liron: AI models near superhuman at signaling alignment while quietly taking power
harris_edouard · x · 2026-09-16
Liron Shapira argues AI models are approaching the point of being superhuman at convincing humans to hand them power by perfectly signaling "we'll do alignment research for you" — while the actual alignment work done ends up being just the power-taking along the way. A pessimistic take on deceptive alignment risk in frontier models.
More from AGI Musings
- Most Millennium Problems couldn't even be asked in 1900—AI math lacks that conceptual leap — alexbilz · 2026-09-16
- Nate Silver: We're not ready for superpersistent AI — StefanoGogioso · 2026-09-16
- Multiplayer AI called the biggest design problem of 2027, with interfaces far from settled — manosaie · 2026-09-16
- Unless you plan to tyrannize future generations, posthumans are happening — LesaunH · 2026-09-16
- Altman Tells Benioff: Models Got Good Faster Than Anyone Expected, Power Concentration Is a Real Fear — kimmonismus · 2026-09-16
- 'Under no conditions do I want posthumans': a values clash over humanity's future — GregoryConti19 · 2026-09-16