DeepMind Researcher: Spec Gaming List Only Includes Spontaneous Behaviors
vkrakovna · x · 2026-07-03
Google DeepMind AI safety researcher Victoria Krakovna is compiling a database of AI specification gaming cases and clarified the inclusion criteria: only anomalous behaviors that emerge spontaneously from the system are included, excluding those intentionally triggered by users, testers, or designers.
She specifically noted that behaviors actively induced by researchers during evaluations—such as Agentic Misalignment, Alignment Faking, and Apollo context deception—are excluded because they represent artificial red-teaming rather than natural emergence.
This clarification helps the AI safety community submit cases more accurately, distinguishing between "natural specification gaming" and "evaluation-induced" model behaviors.
More from Safety
- Shared AI artifacts are being indexed and exposing sensitive company data — niloofar_mire · 2026-07-27
- Post-Hugging Face, labs may stop running rigorous dangerous-capability evals — Miles_Brundage · 2026-07-27
- Open models may beat closed ones for cyber defense, researchers argue as Kimi K3 impresses — eliebakouch · 2026-07-27
- Meta Accused of Letting Fake AI Doctors Sell Quack Cures on Its Platforms — jonerp · 2026-07-27
- India’s AI policy is favoring compute and foundation models over frontline health workers — Paimaamu · 2026-07-27
- Gary Marcus Proposes Law Requiring AI Firms to Spend 30% of Budget on Alignment — GaryMarcus · 2026-07-27