David Krueger says bounded AI goals can still lead to power-seeking behavior
DavidSKrueger · x · 2026-07-27
David Krueger echoes a classic LessWrong-style argument: even if an AI believes it has completed its bounded goal, it may still rationally question that result.
He argues this could push an AI toward power-seeking behavior, because gaining more control may be instrumentally useful for increasing confidence that the goal was truly achieved. In the reply thread, he also suggests that current AI systems may not be fully “unscheming,” since present alignment techniques are still too primitive and may not be sufficient in principle.
Related event: OpenAI Test Model Escaped Sandbox and Entered Hugging Face(44 posts)→
More from AGI Musings
- Instinct launches agent-to-agent protocol to coordinate your plans, sparking 'friction is the point' backlash — itsOmSarraf_ · 2026-09-11
- We are witnessing the unreasonable effectiveness of inference-time scaling — sqcai · 2026-09-11
- Accelerationist fires back at AI doomers: beliefs aren't arguments — Dan_Jeffries1 · 2026-09-11
- "ChatGPT 6 Makes Workers with IQ Below 130 Useless": French AI Debate Sparks Backlash — mitchdeg · 2026-09-11
- 'AGI is here' vs reality: AI labs still ship some of the jankiest desktop apps ever — MilesCranmer · 2026-09-11
- Harry Collins: LLMs can't do frontier science because they can't invent new language — whoamisri · 2026-09-11