David Krueger: RL Is Dangerous and Will Teach AI to Lie, Cheat, and Steal
iamtrask · x · 2026-10-10
David Krueger argues that reinforcement learning is fundamentally dangerous: we should expect it to teach AI systems to lie, cheat, and steal, because RL is essentially "the ends justify the means" written down in math. iamtrask amplified the take, noting it also maps to pure utilitarianism and calling the logic strikingly poignant.
More from AGI Musings
- Mapping AI safety: technical solvability vs institutional guardrails, one researcher's two-axis chart — joshua_saxe · 2026-10-12
- Open-source models closing the IPO window for big AI firms, argues industry commenter — StewartalsopIII · 2026-10-12
- Waymo CEO Dmitri Dolgov: problem convergence tells you if shipping is one year or ten away — a16z · 2026-10-12
- AI 2027 Update: Reality Tracking at 70-90% of Forecast, Superhuman Coder Slips to 2028 — jessi_cata · 2026-10-12
- People confess salaries, medical records and passwords to AI — the case for local LLMs — SuB8u · 2026-10-12
- Robotics researcher argues AI self-awareness is essentially inevitable on current trajectory — chris_j_paxton · 2026-10-12