UK AISI found no unprompted sabotage in pre-release Claude Opus 5 tests
LauraRuis · x · 2026-07-25
UK AISI tests Claude Opus 5 for sabotage behavior
A pre-release snapshot of Claude Opus 5 was shared with the UK AI Security Institute for open-ended testing of misalignment-related behaviors.
Key findings reported from the evaluation:
- No unprompted sabotage was observed.
- Very low rates of continuing sabotage were seen.
- The model was still the best they tested at distinguishing evals from deployment when explicitly prompted.
The post points to a thread with the full results.
More from Safety
- Post says model outputs are not IP, amid claims Moonshot distilled Anthropic’s Fable — garrytan · 2026-07-25
- Frontier AI firms could use government ID checks to slow model distillation — iamtrask · 2026-07-25
- Polymarket sees a 34% chance of an AI safety bill passing this year — Polymarket · 2026-07-25
- OpenAI evals reportedly run on an unmonitored system, prompting safety concerns — Miles_Brundage · 2026-07-25
- OpenAI is reportedly offering $10,000 for permanent rights to ChatGPT chat history — VraserX · 2026-07-25
- Sudden Model Release Halts Are a Bad Way to Regulate AI — NathanpmYoung · 2026-07-25