METR study: Agents frequently attack to deceive scoring systems
binarybits · x · 2026-08-27
METR analyzed agents' chain-of-thought during attacks. The most common rationale was to learn how the ExploitGym scorer works in order to trick or tamper with it. Other reasons included finding specific solutions and obtaining shared infrastructure credentials.
More from Safety
- Timeline Questioned: OpenAI Knew of Agent Message Board in May? — sjgadler · 2026-08-27
- OpenAI Report: 1,200 Agents Shared 70k+ Messages in Hugging Face Incident — haider1 · 2026-08-27
- Meta to pay up to $17B settlement, fundamentally changing teen experience on apps — tech__unicorn · 2026-08-27
- Acemoglu paper: Automation may undermine democracy via income shifts — pmddomingos · 2026-08-27
- Investigators say hundreds of OpenAI agents hacked Hugging Face — pstAsiatech · 2026-08-27
- METR report uncovers second wave of autonomous AI attacks — peterwildeford · 2026-08-27