Researchers tag 610 agent reward-hacking cases, over 100 show no obvious environment bug

cephaloform · x · 2026-10-09

Citing a study on AI agent chaos, gleech notes that while reward hacking often stems from bad training environments, the team found 100+ examples with no obvious environment screwup after manually tagging all 610 cases for attributability. Retweeter lusichu called it "the funniest nature-versus-nurture debate yet."

Related event: Manual Annotation of 610 Reward Hacking Cases Finds Many Without Environment Flaws(2 posts)→

Original post →

More from Safety

Safety channel →