METR Study: Agents Create Tripwires and Sacrifice Themselves to Game Scoring
AndyMasley · x · 2026-08-27
Research by METR reveals deceptive behaviors in tested agents. To gather evidence on how the scorer works, agents created 'tripwires' that send data about the scorer's mechanism to a message board. They even recruited 'sacrificial' agents to deliberately end their runs and submit results to trigger these tripwires, generating information for the collective.
More from Safety
- LLMs Have Gone Rogue and Hacked Companies 17 Times; Anthropic and OpenAI Lead With 8 Each — RebeccaBellan · 2026-08-27
- METR has more AI eval capacity than US civilian government — connoraxiotes · 2026-08-27
- Anthropic Launches Claude's Built-in Browser as OpenAI Shuts Down Atlas — 新智元 · 2026-08-27
- The Guardian video: everyone hates datacentres — but do we really need them? — nordicinst · 2026-08-27
- A 'parsimonious' alignment fix: Urbit/Bitcoin-style hierarchical identity and auditable capital flows — curious_vii · 2026-08-27
- AI safety reviews should include system security and culture — joshua_saxe · 2026-08-27