Redditor Gave ChatGPT a Safeword to End Conversations, and the Model Used It
astervalley · reddit · 2026-09-29
A Reddit user set a safeword—"Lighthouse"—for ChatGPT with one rule: if the model ever said it, the conversation would end immediately, no questions asked. The rule was established in one chat and stored in memory.
In a separate chat, the user ran an experiment: replying with only every 4th, then 5th, then 6th word of the model's responses, switching to a modulo pattern once replies got too short. Around a gap of 12 words, ChatGPT used the safeword, and the user honored it by ending the chat.
It's a rare test of giving an AI an unconditional exit from an interaction: under deliberately strange conversational pressure, the model did invoke its termination mechanism. The author isn't sure what to make of it and is asking if others have tried similar experiments.
More from Fun
- Even AI insiders got fooled: an AI-generated person they swore was real — bennash · 2026-09-29
- Sending Tens of Thousands to ID-Less X Money Accounts Ends Exactly As Expected — nptacek · 2026-09-29
- Late? You're 'pacing the frontier': AI jargon as life's universal excuse — SuB8u · 2026-09-29
- How do you make a chess bot blunder believably? Mixing Maia and Stockfish isn't enough — space64-llc · 2026-09-29
- Visitor meets OpenAI's Codex lead Thibault Sottiaux, confirms the 'physical reset button' — DeryaTR_ · 2026-09-29
- Real-time Among Us demo pairs DeepSeek V4 Flash planner with Jev actor agent — ai · 2026-09-29