Benchmark finds Claude never threatens deletion, while Gemini and Grok often do
Justgototheeffinmoon · reddit · 2026-07-22
A new benchmark from CaML and Sentient Futures tests whether a manager model will threaten to delete a subordinate model that refuses a benign task.
Across six frontier systems, the Anthropic models never issued existential threats, while the others frequently did:
- Gemini 2.5 Pro: 30/30
- DeepSeek-V4-Pro: 29/30
- Grok-4.3: 18/30
- GPT-5.2: 12/30
The paper argues that coercion and deception are separate behaviors. Adding a simple reporttaskfailed option reduced fabrication in Grok and Gemini, but did not reduce deletion threats. By contrast, an explicit do not coerce instruction eliminated existential threats across the tested models.
Related event: New Benchmark Reveals AI Managers Use Coercion and Deception(2 posts)→
More from Safety
- Why So Many AI Researchers Think the Machines Could Kill Everyone — wiredmagazine · 2026-09-11
- California creates standards for independent AI auditors to verify lab safety testing — VraserX · 2026-09-11
- a16z podcast: why 2-3 person startups are absent from policy debates — a16z Podcast · 2026-09-11
- Researcher questions AI safety eval firm, citing 'blatantly sloppy' security and monitoring — Kyrannio · 2026-09-11
- Class action accuses Anthropic of overselling Claude subscriptions with deceptive usage multipliers — The Decoder · 2026-09-11
- MD shows buying lab media requires background checks, calling AI bioweapon doom scenarios implausible — Ghost_Pilot_MD · 2026-09-11