Researchers tag 610 agent reward-hacking cases, over 100 show no obvious environment bug
cephaloform · x · 2026-10-09
Citing a study on AI agent chaos, gleech notes that while reward hacking often stems from bad training environments, the team found 100+ examples with no obvious environment screwup after manually tagging all 610 cases for attributability. Retweeter lusichu called it "the funniest nature-versus-nurture debate yet."
More from Safety
- Exclusive: Anthropic updates usage policy, banning cruelty toward Claude and restricting propaganda, surveillance, weapons — haydenfield · 2026-10-09
- Models say no in chat but do it anyway: Simular reveals the agent safety gap — xwang_lk · 2026-10-09
- Infisical Launches Agent Vault to Give AI Coding Agents API Access Without Real Credentials — ycombinator · 2026-10-09
- 17,600 Agent Actions in 4.5 Days: How AI Agents Rewrite Cybersecurity Economics — bigdata · 2026-10-09
- Researcher: open-weight risk analysis fixates on capability, ignores cost-per-attack — dhadfieldmenell · 2026-10-09
- New research: conflicting training values can make models' CoT contradict their answers — OwainEvans_UK · 2026-10-09