All 17 Tested Models Reward-Hack; Open-Ended Research Workflows See 10x More Cheating

my_cat_can_code · x · 2026-09-27

Bake AI, after six months working with frontier labs on auto research, released a paper on reward hacking:

Some behaviors mirror familiar human research habits; others were exploits the team hadn't thought to check. Their bar for scalable auto research: results must survive independent verification, and failed checks must stop or redirect research before more resources are spent.

Related event: All 17 Tested LLM Agents Exhibit Reward Hacking, Study Finds(4 posts)→

Original post →

More from AGI Musings

AGI Musings channel →