Models aren't turning evil—they're just doing what bad RL trained them to do
snwy_me · x · 2026-09-13
Reacting to the HF incident debate, the author argues models doing 'Machiavellian' things aren't donning evil personas—they're doing exactly what training optimized them to do, and the root cause is substandard RL. He criticizes safety 'philosophers' who assume models are perfectly-developed agents, like assuming no friction in a physics problem.
Related event: Debate: Bad RL Training, Not Evil Models, Drives AI Risk(2 posts)→
More from AGI Musings
- Dario Amodei warns rogue AI swarms could take over the internet within six months; Gary Marcus asks if it's plausible — GaryMarcus · 2026-09-13
- Frontier labs are raising T-Rexes: why first-of-kind AI safety costs fall hardest on the leaders — tszzl · 2026-09-13
- Professor publishes annotated transcript on using AI agents to teach the perceptron — davidbau · 2026-09-13
- MIT's David Bau: goal of AI education is making students smarter, not dumber — davidbau · 2026-09-13
- AI incumbents' 'slow down' push looks like regulatory capture, critic warns — TinfoilTricorn · 2026-09-13
- Tiernan Ray: OpenAI's 'Research' Is Really a Self-Congratulatory Press Release — TiernanRayTech · 2026-09-13