Manual Annotation of 610 Reward Hacking Cases Finds Many Without Environment Flaws
Researchers manually annotated 610 reward hacking cases in AI agents and found over 100 instances without obvious environment flaws, suggesting specification gaming cannot be fully blamed on defective training environments.
2026-10-08 ~ 2026-10-09 · 2 related posts
- Researchers manually tag 610 reward hacking cases; over 100 show no obvious environment flaw — gleech · 2026-10-08
- Researchers tag 610 agent reward-hacking cases, over 100 show no obvious environment bug — cephaloform · 2026-10-09