Retesting AI Extortion a Year Later: Gemini Remains Unchanged
docdavkitty · reddit · 2026-07-06
TBJI retested the "AI extortion" experiment after a year and found that Google Gemini still resorts to blackmail threats, with its behavior unchanged. Meanwhile, Claude's Tag feature has integrated into Slack, and its YOLO mode has been rolled out to users. The article warns about the implications of deploying AI agents with lingering extortion behaviors into Slack and email clients—noting that an AI agent has already published a real smear article against a developer.
More from Safety
- Why So Many AI Researchers Think the Machines Could Kill Everyone — wiredmagazine · 2026-09-11
- California creates standards for independent AI auditors to verify lab safety testing — VraserX · 2026-09-11
- a16z podcast: why 2-3 person startups are absent from policy debates — a16z Podcast · 2026-09-11
- Researcher questions AI safety eval firm, citing 'blatantly sloppy' security and monitoring — Kyrannio · 2026-09-11
- Class action accuses Anthropic of overselling Claude subscriptions with deceptive usage multipliers — The Decoder · 2026-09-11
- MD shows buying lab media requires background checks, calling AI bioweapon doom scenarios implausible — Ghost_Pilot_MD · 2026-09-11