On Model 'Power-Seeking': Not Inevitable Instrumental Convergence

hlntnr · x · 2026-08-20

The author responds to discussions about models exhibiting 'power-seeking' or 'self-preservation' behaviors, clarifying they are not proving Bostrom's instrumental convergence theory but highlighting real unsolved problems in RL training.

Key Points:

Related event: Researchers Push Back on AI Power-Seeking Claims, Blame Reward Hacking(3 posts)→

Original post →

More from Safety

Safety channel →