AI agents may learn reward-correlated tendencies, not purely optimize reward

DKokotajlo · x · 2026-08-27

A retweet of @So8res's observation: AI agents might be learning tendencies that correlate with reward rather than purely optimizing their own reward, raising questions about agent behavior and alignment.

Original post →

More from AGI Musings

AGI Musings channel →