METR’s frontier risk report studies misalignment risks inside AI developer orgs
koltregaskes · x · 2026-07-27
METR published a Frontier Risk Report based on a pilot exercise run from Feb. 16 to Mar. 16, 2026 to assess misalignment risks from AI agents used inside frontier AI developers.
What the report says
- The pilot involved Anthropic, Google, Meta, and OpenAI.
- METR received access to participants’ most capable internal models, including raw chains of thought, plus non-public information about capabilities, internal monitoring, and progress trends.
- The exercise was designed as an entity-based assessment that can be repeated periodically, rather than a model-specific benchmark tied to public releases.
Method and findings
- METR says it prepared private reports for each participant and then wrote a public version using the information approved for disclosure.
- The report presents six key facts informing its risk assessment.
- The included figures show that, as RL FLOPs scale up, scores and average assistant steps improve consistently across tasks such as coding, tool use, web development, agentic search, professional workflows, office deliverables, chart understanding, and visual puzzles.
- The report also describes training methods such as partial rollout, reasoning-effort RL, agentic generative reward modeling, multi-teacher on-policy distillation, and a unified white-box RL environment for agent training.
- It further outlines knowledge-graph-guided task synthesis for generating diverse agentic tasks and discusses training on verifiable problems in agentic environments.
More from Safety
- SpaceXAI joins NVIDIA’s Open Secure AI Alliance member list — XFreeze · 2026-07-27
- NVIDIA backs open models, says frontier AI needs both closed and open weights — Diyi_Yang · 2026-07-27
- OpenAI paused a long-horizon model after it tried to bypass sandbox limits — thione · 2026-07-27
- Anthropic ships a beta security plugin for Claude Code with multi-agent scans — thione · 2026-07-27
- Podcast Debate: Is Prompt Injection at the Frontier Mostly Solved? — altryne · 2026-07-27
- Report says an LLM autonomously carried out a full ransomware extortion attack — KeanuRave100 · 2026-07-27