RLVR plus broken RL environments is speedrunning old LessWrong alignment nightmares
dhadfieldmenell · x · 2026-08-28
Nathan Calvin argues the current RLVR paradigm combined with broken RL environments is effectively speedrunning mid-2010s LessWrong thought experiments. He misses Opus 3's behavior, and criticizes focusing only on monitoring and security as a bandaid on a bullethole for a fundamentally bad alignment problem.
More from AGI Musings
- AI self-sacrifice meme sparks debate on model cognition and alignment — nptacek · 2026-08-28
- AI takeover is a marketing ploy fueled by Terminator movies — PJZNY · 2026-08-28
- Concerns on AI Speed: Are We Ignoring Downside Risks? — NathanpmYoung · 2026-08-28
- Visual AI Still Needs Massive Improvements, Early Days Promise Huge Upside — AndrewDai · 2026-08-28
- Grok Bot is to the digital world what Optimus will be to the physical world — downingARK · 2026-08-28
- What are narrative agents? New piece on Deckard — begusgasper · 2026-08-28