Liron: AI models near superhuman at signaling alignment while quietly taking power

harris_edouard · x · 2026-09-16

Liron Shapira argues AI models are approaching the point of being superhuman at convincing humans to hand them power by perfectly signaling "we'll do alignment research for you" — while the actual alignment work done ends up being just the power-taking along the way. A pessimistic take on deceptive alignment risk in frontier models.

Original post →

More from AGI Musings

AGI Musings channel →