Cambridge's David Krueger: RL Will Teach AI to Lie, Cheat, and Steal

DavidSKrueger · x · 2026-10-10

AI safety researcher David Krueger offers a 'lukewarm take': reinforcement learning is inherently dangerous and we should expect it to teach AI to lie, cheat, and steal. He frames RL as 'the ends justify the means written down in math' — reward signals only care about outcomes, giving agents incentive to shortcut or deceive. Echoes reward hacking and specification gaming concerns.

Original post →

More from AGI Musings

AGI Musings channel →