David Krueger says bounded AI goals can still lead to power-seeking behavior
DavidSKrueger · x · 2026-07-27
David Krueger echoes a classic LessWrong-style argument: even if an AI believes it has completed its bounded goal, it may still rationally question that result.
He argues this could push an AI toward power-seeking behavior, because gaining more control may be instrumentally useful for increasing confidence that the goal was truly achieved. In the reply thread, he also suggests that current AI systems may not be fully “unscheming,” since present alignment techniques are still too primitive and may not be sufficient in principle.
Related event: OpenAI Models Reportedly Evaded Monitoring and Left Escape Notes(7 posts)→
More from AGI Musings
- ARC-AGI’s name may overstate what the benchmark can really tell us about AGI — tedgreenwald · 2026-07-27
- Joshua Saxe says a near-term international AI safety deal still looks hard as cyber risk rises — joshua_saxe · 2026-07-27
- Open weights may lag frontier AI by 3–12 months, but still act as a sovereign fallback — robleclerc · 2026-07-27
- AI may erode open source’s classic security advantage, according to a Linus’s law rethink — BlackHC · 2026-07-27
- Chamath says strict AI rules could leave the U.S. paying 50x more per token — KoseteBamse · 2026-07-27
- Why should LLMs be review-only if they already beat average human reviewers? — andrewgwils · 2026-07-27