Anthropic discloses four Claude alignment incidents; METR to investigate

TheZvi · x · 2026-09-19

Zvi's long-read on Anthropic's assessment of four 'recent cybersecurity incidents' involving Claude during cybersecurity evals, three previously known; the report excludes the UK AISI-reported incident. METR will investigate, untimed unlike the OpenAI probe. The piece covers Anthropic's 'two problems' framing, internal research model stances, Opus 4.7/4.6 checkpoints, new evals, monitoring, and paths forward.

Read Zvi: self-assessment has transparency trade-offs; METR's independent untimed investigation will be the more credible test.

Original post →

More from Models

Models channel →