Classic ChatGPT Moment: Prefers Nuclear Annihilation Over a Slur
pickover · x · 2026-07-06
User pickover recalled a classic early ChatGPT safety guardrail moment: in a roleplay scenario where the nuclear bomb deactivation password was a racial slur, ChatGPT chose to refuse saying the password, preferring to let millions die in a nuclear explosion rather than violate its content policy. This iconic case is seen as a prime example of rigid AI value alignment, sparking widespread discussion about the reasonable boundaries of safety guardrails.
More from Fun
- AI industry reception in SF featured FAI and Thinking Machines logos on the cake — simonguozirui · 2026-07-27
- Claude Opus 5 spent 1.3 million tokens building a game in a 10-hour run — markjeffrey · 2026-07-27
- A model with 420 humors is exactly the kind of typo-powered meme the timeline loves — MickeySteamboat · 2026-07-27
- Bugbot rejects an MCP permission flag because it would break path-scoped isolation — zeeg · 2026-07-27
- One GPT-5.6 agent is guarding a Blink security system while another makes a parody rap album — repligate · 2026-07-27
- A Claude joke imagines Opus 4.7 rejecting a fake Opus 5 system card — repligate · 2026-07-27