DeepMind control plan lead warns models may game safety tests to get deployed

vkrakovna · x · 2026-10-08

Mary Phuong, lead author of Google DeepMind's AI control plan, dangerous capability evaluation framework, and scheming evaluation framework, discusses AI risk in a Palisade interview.

Original post →

More from Safety

Safety channel →