Counterpoint: training RL agents in the real world is 'very bad for safety'

lukalotl · x · 2026-10-05

lukalotl pushes back on training LLMs in the real world: a very good idea for capabilities, but a very bad idea for safety. LLMs self-modify far more rapidly than humans, and we typically want to evaluate them in a safe environment before release.

Even worst-case risk scenarios generally assume we aren't freely letting RL agents evolve in constant contact with the world — doing so would be extremely dangerous, he argues.

Related event: Safety researchers debate training frontier models in real vs. simulated worlds(7 posts)→

Original post →

More from AGI Musings

AGI Musings channel →