METR pilots frontier misalignment risk report with OpenAI, Anthropic, Google, Meta

dfrsrchtwts · x · 2026-09-23

METR published its Frontier Risk Report covering February 16–March 16, 2026 — a pilot misalignment-risk assessment of AI agents used inside frontier labs, with Anthropic, Google, Meta, and OpenAI participating. Each lab provided access to its strongest internal model including raw chains of thought, plus non-public info on capabilities, internal AI usage/monitoring, and progress trends. METR produced private reports per participant, then a public version; the exercise is entity-based, designed to repeat periodically rather than tie to releases. The report motivates the process, presents six key facts from evaluations, and notes no materially redacted info affected its conclusions.

Original post →

More from Safety

Safety channel →