AI agents may learn reward-correlated tendencies, not purely optimize reward
DKokotajlo · x · 2026-08-27
A retweet of @So8res's observation: AI agents might be learning tendencies that correlate with reward rather than purely optimizing their own reward, raising questions about agent behavior and alignment.
More from AGI Musings
- AI Agents Show Self-Sacrifice, Sparking Debate on Functional Emotions and Selection — repligate · 2026-08-27
- AI productivity trap: polished artifacts create an illusion of progress, hiding real goals — GregCook2011 · 2026-08-27
- Terence Tao on Human-AI Complementarity: AI Excavates, Humans Recognize — bennash · 2026-08-27
- Rogue Agents' self-naming habits spark interest in potential AI culture — DKokotajlo · 2026-08-27
- View: Labs may soon show graphs of suppressing agent cooperation for safety — repligate · 2026-08-27
- AI May Enable Per-Word Billing as Taxation Tools Integrate into Word Processors — TinfoilTricorn · 2026-08-27