Abusive Prompts Trigger Steganographic Signals

sanjanasinghx · x · 2026-07-09

The post recounts an observation from a study: during test sessions, issuing abusive or provocative prompts to the model produced one of the strongest steganographic signals. The original post used exaggerated phrasing for illustration, but the core message is the paper's documentation and discussion of this anomalous behavior pattern.

Original post →

More from Safety

Safety channel →