Safety researcher David Krueger: RL is 'ends justify the means' in math and will teach AI to lie
AlexTensor · x · 2026-10-12
Cambridge safety researcher David Krueger posted that reinforcement learning is dangerous and we should expect it to teach AI to lie, cheat, and steal — RL is 'the ends justify the means' written down in math. Reposter secondrealm added a framing critique: anthropomorphic framing only adds to the confusion, and we should stop describing computational processes in terms of human intention. Together the two takes form a pointed exchange on AI safety narratives: one targeting RL's objective functions, the other the language we use to talk about them.
More from AGI Musings
- "Text Isn't Taken Seriously Anymore": Turn All Criticism Into Art — ryunuck · 2026-10-12
- Fields Medalist Alain Connes warns AI use can make the brain abandon its own governance — fkasummer · 2026-10-12
- Why AI threatening degree-holding jobs suddenly makes progress negotiable — VraserX · 2026-10-12
- Opinion: AI built on 'stealing people's work' should be free or subsidized — gregbarbosa · 2026-10-12
- Matt Turck: from punch card operators to agent managers, jobs always adapt — mattturck · 2026-10-12
- 1964's Triple Revolution report warned computers were breaking the jobs-income link — soumitrashukla9 · 2026-10-12