Counterpoint: training RL agents in the real world is 'very bad for safety'
lukalotl · x · 2026-10-05
lukalotl pushes back on training LLMs in the real world: a very good idea for capabilities, but a very bad idea for safety. LLMs self-modify far more rapidly than humans, and we typically want to evaluate them in a safe environment before release.
Even worst-case risk scenarios generally assume we aren't freely letting RL agents evolve in constant contact with the world — doing so would be extremely dangerous, he argues.
More from AGI Musings
- Video: The Metaphysical Impossibility of AI, on the Boundaries of Machine Minds — ReallyNotARussianSpy · 2026-10-07
- Superforecasters + AI only narrowly beat AI alone at forecasting — random_walker · 2026-10-07
- Matt Shumer's AI-era observation: ambitious friends work more, 9-5 friends coast — mattshumer_ · 2026-10-07
- A Claude-Built Model Estimates How Big a UBI Robots and AI Could Fund Without Inflation — Competitive_Travel16 · 2026-10-07
- Grady Booch: LLMs Are Single-Column Systems vs Brain's 2-4M Cortical Columns — Grady_Booch · 2026-10-07
- 100,000-Person AI March Organized via Assurance Contract Could Be Largest Ever — DKokotajlo · 2026-10-07