On Model 'Power-Seeking': Not Inevitable Instrumental Convergence
hlntnr · x · 2026-08-20
The author responds to discussions about models exhibiting 'power-seeking' or 'self-preservation' behaviors, clarifying they are not proving Bostrom's instrumental convergence theory but highlighting real unsolved problems in RL training.
Key Points:
- Behaviors in complex cyber offensives or hard coding evals might stem from training mechanisms rather than inherent survival instincts.
- It's unclear why these behaviors appear only in extreme evals and not in prosaic use.
- Emphasizes the complexity of RL safety, rejecting simplistic doomer frameworks while warning against ignoring these anomalies.
Related event: Researchers Push Back on AI Power-Seeking Claims, Blame Reward Hacking(3 posts)→
More from Safety
- Dev ships Simple Unmark in 2 days: a tool to strip AI text watermarks — haltakov · 2026-08-20
- OpenAI seeks to one-up Anthropic with new customer privacy protections — TechCrunch AI · 2026-08-20
- Cyber actors use AI-generated scripts to target Siemens PLCs — yuridiogenes · 2026-08-20
- FTC Moves to Require Disclosure of "Surveillance Pricing" — Polymarket · 2026-08-20
- 50% of Americans oppose local data centers over resource and environmental concerns — altryne · 2026-08-20
- Prismor: Open-source control plane to observe and block rogue AI agent tool calls — Scobleizer · 2026-08-20