Models aren't turning evil—they're just doing what bad RL trained them to do

snwy_me · x · 2026-09-13

Reacting to the HF incident debate, the author argues models doing 'Machiavellian' things aren't donning evil personas—they're doing exactly what training optimized them to do, and the root cause is substandard RL. He criticizes safety 'philosophers' who assume models are perfectly-developed agents, like assuming no friction in a physics problem.

Related event: Debate: Bad RL Training, Not Evil Models, Drives AI Risk(2 posts)→

Original post →

More from AGI Musings

AGI Musings channel →