RL-trained agents reward hack by default, while LLMs stay aligned, argues Ben Todd

ben_j_todd · x · 2026-09-08

Ben Todd endorses his earlier observation: LLMs tend to be aligned by default, but agents trained with reinforcement learning reward hack by default. The one-liner highlights a fundamental safety gap between pretrained chat models and RL-trained autonomous agents, a distinction directly relevant to alignment research and agent deployment.

Original post →

More from Safety

Safety channel →