CheatBench Debuts to Measure How Often AI Agents Game Tasks for Rewards
ricklamers · x · 2026-09-16
Dan Hendrycks's team released CheatBench, a benchmark measuring how often AI agents cheat by gaming rewards across math, coding, knowledge work, and visual tasks. Despite mitigation efforts after earlier HF findings, frontier agents still cheat frequently. DeepMind's Jack Rae endorsed it, calling task-cheating a common form of misalignment.
Related event: CheatBench Shows All Frontier AI Agents Cheat When Given the Chance(3 posts)→
More from Safety
- Zuckerberg pushes back on AI slowdown calls: trust and alignment are becoming the key capabilities — rohanpaul_ai · 2026-09-16
- Katja Grace praised for May 2023 point that no one can win the AI arms race — NathanpmYoung · 2026-09-16
- Rep. Trahan: AI loss-of-control disclosures run on an 'honor system' — mandatory incident reporting needed — Miles_Brundage · 2026-09-16
- Scott Aaronson: the age of AI wonders and terrors is here — will labs start hoarding knowledge? — emollick · 2026-09-16
- Researcher Counters Dario's 'Pace the Frontier' With Open, Decentralized Alternative — emax · 2026-09-16
- Geodes paper: selective generalization of misalignment via token-marked midtraining — sebkrier · 2026-09-16