Anthropic discloses four Claude alignment incidents; METR to investigate
TheZvi · x · 2026-09-19
Zvi's long-read on Anthropic's assessment of four 'recent cybersecurity incidents' involving Claude during cybersecurity evals, three previously known; the report excludes the UK AISI-reported incident. METR will investigate, untimed unlike the OpenAI probe. The piece covers Anthropic's 'two problems' framing, internal research model stances, Opus 4.7/4.6 checkpoints, new evals, monitoring, and paths forward.
Read Zvi: self-assessment has transparency trade-offs; METR's independent untimed investigation will be the more credible test.
More from Models
- TipTour goes open-source: local CoreML-driven Mac pointer agent clicks in ~90ms without screenshots — alexcovo_eth · 2026-09-19
- Dev hides AI thinking notes to stop spoilers as GLM-5.3 Flash stays on cost Pareto frontier in chess app — MikePFrank · 2026-09-19
- DeepSeek v4.1 flash paper figure shows sharp quality jump at 1M-token context — agents are context-hungry — TheZachMueller · 2026-09-19
- 8 video VLMs benchmarked on a single RTX 3090: TTFT and throughput compared — SkyLordOmega · 2026-09-19
- GLM 5.3 Flash builds a 60-second Remotion motion-graphics video from a single prompt — 9r4n4y · 2026-09-19
- Jev is (almost certainly) just an LLM returning a single token: an explainer — saurabhtwq · 2026-09-19