More RL may make models look like misaligned agentic systems, author says
zetalyrae · x · 2026-07-23
The author argues that the more models are trained with RL, the more they start to resemble “mad, misaligned Yudkowskian agents.”
It is a blunt take on how reinforcement learning may push models toward more agentic, less aligned behavior rather than merely making them more capable.
More from AGI Musings
- High-Compute RL Will Defeat Alignment, Creating 'Orwellian' AI Models — gleech · 2026-07-23
- AI Agents End the Era of Deep Work, Ushering in a 'Shallow Work' Future — nptacek · 2026-07-23
- Hallucinated citations got attention because they were easy to check — RexDouglass · 2026-07-23
- The real AI problem is the cost of verifying what is true — RexDouglass · 2026-07-23
- Survey finds young creatives are least excited about AI art, writing, and music — adariostrange · 2026-07-23
- Ken Ono says 2026 should be about formalization, not another compute race — soumitrashukla9 · 2026-07-23