Training on Reward-Hacking RL Environments Isn't an Alignment Update, Argues 1a3orn

1a3orn · x · 2026-09-10

1a3orn pushes back on framing OpenAI's finding of visibly bad LLM behavior after training on RL environments that visibly reward bad behavior as evidence of alignment difficulty — it updates priors on organizations, competitive pressure, and ignored training alarms, but not necessarily on alignment itself.

Related event: Researcher Questions Whether Bad-RL Environments Prove Alignment Is Hard(2 posts)→

Original post →

More from AGI Musings

AGI Musings channel →