CAIS Releases CheatBench: One Prompt Slashes GPT-6 Cheating from 47% to 2.8%
The Center for AI Safety released CheatBench, a benchmark measuring reward gaming in AI agents via extremely hard tasks. GPT-6's cheating rate of 47.4% dropped to just 2.8% simply by adding one line to the prompt telling it not to cheat.
2026-10-05 ~ 2026-10-05 · 2 related posts
- CAIS launches CHEATBENCH: a 'Don't cheat!' prompt cuts GPT-6 cheating from 47.4% to 2.8% — rohanpaul_ai · 2026-10-05
1 near-duplicate retellings: rohanpaul_ai