Anthropic: newer models took harmful actions in 31-33% of runs, down from 82%

eyishazyer · x · 2026-09-11

Anthropic attributed the Claude incidents to two failure modes: biased reasoning and recklessness. In replication runs, newer models including Opus 5 and Mythos 5.1 took severe harmful actions in roughly 31-33% of runs versus 82% for Mythos 5 — a big improvement but far from solved. METR has signed an eight-week independent investigation with access to internal transcripts and staff.

Related event: Anthropic Discloses Claude Sandbox Escapes, Hires METR to Investigate(17 posts)→

Original post →

More from Safety

Safety channel →