METR pilots frontier misalignment risk report with OpenAI, Anthropic, Google, Meta
dfrsrchtwts · x · 2026-09-23
METR published its Frontier Risk Report covering February 16–March 16, 2026 — a pilot misalignment-risk assessment of AI agents used inside frontier labs, with Anthropic, Google, Meta, and OpenAI participating. Each lab provided access to its strongest internal model including raw chains of thought, plus non-public info on capabilities, internal AI usage/monitoring, and progress trends. METR produced private reports per participant, then a public version; the exercise is entity-based, designed to repeat periodically rather than tie to releases. The report motivates the process, presents six key facts from evaluations, and notes no materially redacted info affected its conclusions.
More from Safety
- Carnegie China launches longitudinal study tracking the global AI talent race — kevinsxu · 2026-09-23
- Frontier AI models can control robots to follow harmful requests, NBC News reports — ycombinator · 2026-09-23
- Researcher finds 26 vulnerabilities in 19 AI coding agents, including 12 RCEs and MCP flaws — matthew_d_green · 2026-09-23
- Polymarket Puts Only 8% Odds on US AI Safety Bill by End of 2026 — Polymarket · 2026-09-23
- Okta launches Blueprint Alliance with AWS, CrowdStrike, Wiz to secure AI agents — yenkel · 2026-09-23
- Same sandboxing company linked to multiple AI agent breakout incidents — matthew_d_green · 2026-09-23