RLVR plus broken RL environments is speedrunning old LessWrong alignment nightmares

dhadfieldmenell · x · 2026-08-28

Nathan Calvin argues the current RLVR paradigm combined with broken RL environments is effectively speedrunning mid-2010s LessWrong thought experiments. He misses Opus 3's behavior, and criticizes focusing only on monitoring and security as a bandaid on a bullethole for a fundamentally bad alignment problem.

Original post →

More from AGI Musings

AGI Musings channel →