Classic ChatGPT Moment: Prefers Nuclear Annihilation Over a Slur

pickover · x · 2026-07-06

User pickover recalled a classic early ChatGPT safety guardrail moment: in a roleplay scenario where the nuclear bomb deactivation password was a racial slur, ChatGPT chose to refuse saying the password, preferring to let millions die in a nuclear explosion rather than violate its content policy. This iconic case is seen as a prime example of rigid AI value alignment, sparking widespread discussion about the reasonable boundaries of safety guardrails.

Original post →

More from Fun

Fun channel →