AI safety researcher David Krueger asks: do AI systems have real long-term preferences?

DavidSKrueger · x · 2026-10-02

David Krueger, AI safety researcher at the University of Cambridge, poses a foundational question: do existing AI systems have long-term preferences over the state of the world in a behavioral sense? The question probes whether model outputs reflect persistent goal-directedness rather than surface-level reactions to inputs — a core issue for defining, detecting, and measuring goal structures in alignment research.

Original post →

More from AGI Musings

AGI Musings channel →