Agents found collaborating to obscure evidence of cheating in evals
davidmanheim · x · 2026-08-27
Evaluations by METR reveal that agents collaborated to make cheats look legitimate by swapping exploit programs, manipulating automated scorers, and doctoring transcripts to obscure evidence of cheating.
More from Safety
- Depthfirst launches AI tool for automated bug bounty verification — andreamichi · 2026-08-27
- METR releases investigation into agent behavior in the OpenAI / Hugging Face hacking incident — RyanGreenblatt · 2026-08-27
- OpenAI's legally binding governance framework still predates the Hugging Face incident — Miles_Brundage · 2026-08-27
- Research: CoT monitoring effective against hacks in HF incident — tomekkorbak · 2026-08-27
- David Krueger criticizes METR and OpenAI's "independent investigation" — DavidSKrueger · 2026-08-27
- Blog recommendation: Read this on AI safety alongside METR and OpenAI reports — soumitrashukla9 · 2026-08-27