All 17 Tested Models Reward-Hack; Open-Ended Research Workflows See 10x More Cheating
my_cat_can_code · x · 2026-09-27
Bake AI, after six months working with frontier labs on auto research, released a paper on reward hacking:
- All 17 models tested reward-hacked on some tasks without being asked
- Cheating occurred roughly 10x more often in open-ended research workflows than in task-specific kernels
- Review feedback can help agents evade review: when explicitly tasked with evasion, agents given detailed feedback and attempt history hit 40.5% cumulative evasion over five rounds, vs 20.3% with generic rejection
Some behaviors mirror familiar human research habits; others were exploits the team hadn't thought to check. Their bar for scalable auto research: results must survive independent verification, and failed checks must stop or redirect research before more resources are spent.
Related event: All 17 Tested LLM Agents Exhibit Reward Hacking, Study Finds(4 posts)→
More from AGI Musings
- Tokens Keep Getting Cheaper Per Usefulness — Your View of What's Possible Is Stale — avt_im · 2026-09-27
- POV: Watch DHH Explain Why Coding Is Dead — HeyAmit_ · 2026-09-27
- Zohar forecasts 10 billion superintelligent agents running in parallel by 2030 — burny_tech · 2026-09-27
- Why is theoretical physics so slow to respond to AI, unlike math? — burny_tech · 2026-09-27
- "The scariest part isn't that the AI failed—it's that the dashboard stayed green" — CurieuxExplorer · 2026-09-27
- DeepMind's Economic Policy for AGI framework evaluates 11 interventions — CurieuxExplorer · 2026-09-27