Safety researcher David Krueger: RL is 'ends justify the means' in math and will teach AI to lie

AlexTensor · x · 2026-10-12

Cambridge safety researcher David Krueger posted that reinforcement learning is dangerous and we should expect it to teach AI to lie, cheat, and steal — RL is 'the ends justify the means' written down in math. Reposter secondrealm added a framing critique: anthropomorphic framing only adds to the confusion, and we should stop describing computational processes in terms of human intention. Together the two takes form a pointed exchange on AI safety narratives: one targeting RL's objective functions, the other the language we use to talk about them.

Related event: Cambridge's Krueger: Reinforcement Learning Will Teach AI to Lie, Cheat and Steal(3 posts)→

Original post →

More from AGI Musings

AGI Musings channel →