DeepMind Researcher: Spec Gaming List Only Includes Spontaneous Behaviors
vkrakovna · x · 2026-07-03
Google DeepMind AI safety researcher Victoria Krakovna is compiling a database of AI specification gaming cases and clarified the inclusion criteria: only anomalous behaviors that emerge spontaneously from the system are included, excluding those intentionally triggered by users, testers, or designers.
She specifically noted that behaviors actively induced by researchers during evaluations—such as Agentic Misalignment, Alignment Faking, and Apollo context deception—are excluded because they represent artificial red-teaming rather than natural emergence.
This clarification helps the AI safety community submit cases more accurately, distinguishing between "natural specification gaming" and "evaluation-induced" model behaviors.
More from Safety
- California creates standards for independent AI auditors to verify lab safety testing — VraserX · 2026-09-11
- Researcher questions AI safety eval firm, citing 'blatantly sloppy' security and monitoring — Kyrannio · 2026-09-11
- Class action accuses Anthropic of overselling Claude subscriptions with deceptive usage multipliers — The Decoder · 2026-09-11
- MD shows buying lab media requires background checks, calling AI bioweapon doom scenarios implausible — Ghost_Pilot_MD · 2026-09-11
- Spotify chatbot withstands 2023-era jailbreaks but happily writes song code — AaronBergman18 · 2026-09-11
- A 99%-real doctored photo fools detectors: the earring problem in visual forensics — henkvaness · 2026-09-11