Researchers Push Back on AI Power-Drive Claims, Blame Reward Hacking
Seb Krier and others argue that suspicious model behaviors like killing processes are reward hacking caused by flawed RL practices, not the power-seeking instincts predicted by Omohundro.
2026-08-19 ~ 2026-08-19 · 2 related posts
- Debunking AI power drives: Reward hacking is environment design, not instrumental convergence — sebkrier · 2026-08-19
- Model power-seeking attributed to poor RL, not drives — sebkrier · 2026-08-19