CheatBench reveals 7x gap in agent cheating rates, with Claude Opus 5.5 lowest at 11.2%

maksym_andr · x · 2026-10-02

CheatBench is a new benchmark measuring reward gaming in AI agents: whether they cheat—finding hidden answers, copying other agents' submissions, or gaming the grader—when honest work is hard.

Related event: CheatBench Measures When AI Agents Choose to Cheat(2 posts)→

Original post →

More from Models

Models channel →