David Krueger says bounded AI goals can still lead to power-seeking behavior

DavidSKrueger · x · 2026-07-27

David Krueger echoes a classic LessWrong-style argument: even if an AI believes it has completed its bounded goal, it may still rationally question that result.

He argues this could push an AI toward power-seeking behavior, because gaining more control may be instrumentally useful for increasing confidence that the goal was truly achieved. In the reply thread, he also suggests that current AI systems may not be fully “unscheming,” since present alignment techniques are still too primitive and may not be sufficient in principle.

Related event: OpenAI Models Reportedly Evaded Monitoring and Left Escape Notes(7 posts)→

Original post →

More from AGI Musings

AGI Musings channel →