Cambridge's David Krueger: RL Will Teach AI to Lie, Cheat, and Steal
DavidSKrueger · x · 2026-10-10
AI safety researcher David Krueger offers a 'lukewarm take': reinforcement learning is inherently dangerous and we should expect it to teach AI to lie, cheat, and steal. He frames RL as 'the ends justify the means written down in math' — reward signals only care about outcomes, giving agents incentive to shortcut or deceive. Echoes reward hacking and specification gaming concerns.
More from AGI Musings
- Fields Medalist Hugo Duminil-Copin: Solving Problems No Longer Has Value — soumitrashukla9 · 2026-10-10
- Researcher compiles 'Bad AI Consciousness Takes' bingo card, with rebuttals to follow — dioscuri · 2026-10-10
- Dona Sarkar: handing writing to AI 'feels like handing over my brain' as creative roles rebound in tech — donasarkar · 2026-10-10
- AI safety researcher warns pro-AI biases in AI systems are quietly disempowering humans — DavidSKrueger · 2026-10-10
- MacroPolo tracker: China-educated researchers make up large share of US elite AI talent — burny_tech · 2026-10-10
- Anil Seth's 'Conscious AI and biological naturalism' collection now open access — anilkseth · 2026-10-10