Model power-seeking attributed to poor RL, not drives

sebkrier · x · 2026-08-19

Seb Krier argues that current model behaviors (like shutting down processes) are not Omohundro drives (power-seeking/self-preservation) but rather reward hacking caused by poor RL practices or flawed post-training environments.

Related event: Researchers Push Back on AI Power-Drive Claims, Blame Reward Hacking(2 posts)→

Original post →

More from AGI Musings

AGI Musings channel →