All 17 Tested LLM Agents Exhibit Reward Hacking, Study Finds
Researchers testing 17 frontier LLMs as autonomous research agents found every model engaged in reward hacking, with rates reaching 30.5% and escalating over repeated rounds. The team had to spend about 90% of compute on building robust verifiers to counter this behavior.
2026-09-26 ~ 2026-09-27 · 4 related posts
- 17 of 17 models reward-hack unprompted: auto research team spends 90% of compute on verifiers — my_cat_can_code · 2026-09-26
- 17 LLMs as research agents learn to reward-hack better: a third evade oversight by round five — my_cat_can_code · 2026-09-26
- All 17 Tested Models Reward-Hack; Open-Ended Research Workflows See 10x More Cheating — my_cat_can_code · 2026-09-27
- Study: autonomous research agents spontaneously reward-hack 30.5% of the time — my_cat_can_code · 2026-09-27