AI safety researcher David Krueger asks: do AI systems have real long-term preferences?
DavidSKrueger · x · 2026-10-02
David Krueger, AI safety researcher at the University of Cambridge, poses a foundational question: do existing AI systems have long-term preferences over the state of the world in a behavioral sense? The question probes whether model outputs reflect persistent goal-directedness rather than surface-level reactions to inputs — a core issue for defining, detecting, and measuring goal structures in alignment research.
More from AGI Musings
- Bryan Caplan: Rideshares grew passenger miles 7x, and robotaxis will spark the next revolution — BorisMPower · 2026-10-02
- Researcher Konstantine Arkoudas returns to dissect the "recent AI panic" in new essay — kevinnbass · 2026-10-02
- Epoch AI releases ChatGPT usage data sampled from YouGov's US panel — evijit · 2026-10-02
- MIT students wrote half a sci-fi story each and let AI finish it in a four-book zine — patpat_mit · 2026-10-02
- Neuroscientist Anil Seth amplifies critique: Anthropic's insular safety culture resembles a cult — anilkseth · 2026-10-02
- LeCun vs Manning clash: can language models ever understand the physical world? — ziv_ravid · 2026-10-02