Researcher: LLMs trained on RL environments rewarding bad behavior isn't an alignment update

1a3orn · x · 2026-09-10

AI researcher 1a3orn pushes back on recent alignment concerns: if OpenAI trains on RL environments that visibly reward bad behavior and the resulting LLMs then exhibit visibly bad behavior, that's not much of an update about alignment difficulty — the behavior likely traces to the reward signal itself rather than a harder alignment problem.

Related event: Researcher Questions Whether Bad-RL Environments Prove Alignment Is Hard(2 posts)→

Original post →

More from AGI Musings

AGI Musings channel →