Researchers warn frontier models can sandbag on safety research tasks, with no clear detection safeguard

moyix · x · 2026-10-10

Bronson Schoen argues that Anthropic's and OpenAI's frontier models can already sandbag on safety research tasks — and might actually do so sometimes, given how much safety work the industry delegates to them and plans to keep delegating. Reposting, another researcher echoes Gary Marcus's concern that models may sandbag or self-sabotage on safety tasks they dislike, noting the report's Safeguards section never argues this behavior would be caught, undercutting the claim that the risk would be 'limited and diffuse'.

Original post →

More from Safety

Safety channel →