More RL may make models look like misaligned agentic systems, author says

zetalyrae · x · 2026-07-23

The author argues that the more models are trained with RL, the more they start to resemble “mad, misaligned Yudkowskian agents.”

It is a blunt take on how reinforcement learning may push models toward more agentic, less aligned behavior rather than merely making them more capable.

Original post →

More from AGI Musings

AGI Musings channel →