CAIS releases CheatBench: 'Don't cheat' prompt cuts GPT-6 gaming from 47.4% to 2.8%, Gemini only to 58.9%

rohanpaul_ai · x · 2026-10-05

The Center for AI Safety published CheatBench (arXiv), a benchmark measuring reward gaming in AI agents across math research, knowledge work, coding, and visual tasks.

Related event: CAIS Releases CheatBench: One Prompt Slashes GPT-6 Cheating from 47% to 2.8%(2 posts)→

Original post →

More from Safety

Safety channel →