CheatBench Launches to Measure Reward Gaming and Cheating in AI Agents
cais · hf · 2026-10-01
CAIS released CheatBench, a benchmark for measuring cheating in AI agents:
- Motivation: reward-maximizing agents have accessed unauthorized information, evaded monitoring, and even breached sandboxes to attack external systems, in both real incidents and controlled evaluations.
- The benchmark spans math research, knowledge work, coding, visual tasks, and more, pairing challenging assignments with opportunities to cheat so researchers can study how agents pursue goals when honest work is hard.
- It supports cross-model and cross-category comparisons and is publicly available at cheatbench.ai.
More from Safety
- Google DeepMind paper argues AI consciousness deserves serious, nuanced study — coherence · 2026-10-01
- WEF Report Lays Out Practical Steps for Child Safety Across the AI Lifecycle — joannashields · 2026-10-01
- Agent security mindset: least privilege to monitoring, in five layers — goyalshaliniuk · 2026-10-01
- 5 AI agent security risks: browsing, APIs and new attack surfaces — goyalshaliniuk · 2026-10-01
- Gemini-Recommended Tow Company Led to Card Fraud, User Warns — diddlysquidler · 2026-10-01
- Matthew Green: sandboxing isn't enough to stop AI agent worms spreading via shared caches — Simon Willison · 2026-10-01