METR to Independently Investigate Anthropic's Agent Incidents and Alignment
yuntiandeng · x · 2026-09-14
Eval org METR announced an agreement with Anthropic to independently investigate agent incidents at the company and its models' alignment properties, with plans to publish reports on findings and terms of engagement. Developers responded that publishing raw agent traces openly would let the open-source ecosystem fix issues faster than any single entity.
Related event: METR to Independently Investigate Anthropic's Agent Incidents and Alignment(4 posts)→
More from Safety
- Security researcher lays out 11-step path to a self-replicating AI botnet stealing API keys — joshua_saxe · 2026-09-14
- Pay-for-a-Paper? Fastcode AI packs a $100K contract into a conference publication — maier_ak · 2026-09-14
- AI safety researcher lays out how a self-replicating agent botnet could hijack inference providers — joshua_saxe · 2026-09-14
- GPT-6 Astra: 1M-token context, 2.5x price, and a 'Critical' cyber risk rating — Thirumalaivasan_GJ · 2026-09-14
- Sam Altman lays out the two ways AI progress could go badly wrong: losing control and power concentration — sama · 2026-09-14
- Open Letter to Amodei and Altman Challenges Frontier AI Slowdown Calls: State Keeps Moving — AryHHAry · 2026-09-14