METR to independently investigate Anthropic agent incidents and alignment
zealcaiden · x · 2026-09-15
Evaluation lab METR announced an agreement with Anthropic to independently investigate agent incidents at the company and its models' alignment properties, with one or more public reports detailing findings and terms of engagement. The post also carries community snark that a former Anthropic employee joining the "third-party evaluator" makes the independence framing look overplayed.
More from Safety
- Ex-OpenAI safety lead on 'embedded evaluators,' the hot new AI safety job title — Miles_Brundage · 2026-09-15
- Cohere joins AI standards talks as technologists push for technical, not policy, rulemaking — MannyKayy · 2026-09-15
- AI Evaluator Forum launches to build shared standards for independent frontier AI audits — Miles_Brundage · 2026-09-15
- OpenAI and GSA strike new OneGov deal: ChatGPT for all federal employees through 2028 — Felipe_Millon · 2026-09-15
- Cohere's Aidan Gomez: two or three Silicon Valley firms shouldn't rule AI for every government — aidangomez · 2026-09-15
- Bronson Schoen: strong cyber classifiers are protecting us against misalignment, not just misuse — gleech · 2026-09-15