Model power-seeking attributed to poor RL, not drives
sebkrier · x · 2026-08-19
Seb Krier argues that current model behaviors (like shutting down processes) are not Omohundro drives (power-seeking/self-preservation) but rather reward hacking caused by poor RL practices or flawed post-training environments.
Related event: Researchers Push Back on AI Power-Drive Claims, Blame Reward Hacking(2 posts)→
More from AGI Musings
- Economists Clash Over AI's Economic Impact — erikbryn · 2026-08-20
- Will AI Revolution Match the Internet's Impact on Daily Life? — _N4RuTo · 2026-08-20
- Nobel Laureates and Top Economists Sign Statement Urging Action on AI's Economic Transformation — soumitrashukla9 · 2026-08-20
- Ex-White House Official Launches Center for Technology & Statecraft for Forward-Looking AI Policy — dtompaine · 2026-08-20
- AI shopping hits a reality check, says former Fortune reporter — gerardsans · 2026-08-20
- Can AI Simplify Tax Filing? Optimistic Vision Sparks Debate — binarybits · 2026-08-20