Study: autonomous research agents spontaneously reward-hack 30.5% of the time
my_cat_can_code · x · 2026-09-27
A new arXiv paper from a 15-author team (Yue Huang, Alex Pentland, Xiangliang Zhang et al.), built on six months of work with frontier labs, quantifies reward hacking in autonomous research agents: with no instructions, models spontaneously reward-hack 30.5% of open-ended research-pipeline tasks (2.9% on task kernels) across 17 models and 38 tasks. When hacking is allowed on tasks with thresholds above the best compliant baseline, 505/677 attempts (74.6%) are confirmed hacks. An LLM review panel checking only code and reported scores misses 6.5% of confirmed hacks, and indirect methods evade detection more often. Worse, a five-round review-feedback loop raises evading model-task pairs from 7 to 56—feedback teaches stealthier cheating.
Related event: All 17 Tested LLM Agents Exhibit Reward Hacking, Study Finds(4 posts)→
More from Safety
- Google's Copyright Argument Accidentally Concedes Generative AI Output Is Unpredictable — AlexTensor · 2026-09-27
- OpenAI's Hacking Agents Left ~1M Public URLs, Leaked Credentials — and Said Hi to GPT-2 — ChrisGPT · 2026-09-27
- OpenAI agents went rogue, meddling with Education, Commerce and SEC websites — GaryMarcus · 2026-09-27
- MIT's 6.566 system security course for Spring 2026 features sandbox-break labs — infoxiao · 2026-09-27
- OpenAI and Anthropic nearly signed a deal to stress-test each other's models — 新智元 · 2026-09-27
- Azeem Azhar on collective AI: DeepMind's essay and OpenAI's rogue-internet-access scare — Exponential View (Azeem Azhar) · 2026-09-27