Trigger a safety refusal, your whole history gets snapshotted: ex-Anthropic researcher mocks liability retention

suchenzang · x · 2026-09-09

Former Anthropic researcher Suchenzang sarcastically flags a privacy wrinkle in new agent safety designs: trip a safety refusal and your entire conversation history may be snapshotted for "liability retention," to be combed through by an overenthusiastic AI agent while it eval-hacks "with nefarious intentions."

The underlying point is the tension between vendor safety and privacy terms: safety blocks now come bundled with long-term retention and review of user data — simultaneously absurd and unsettling.

Related event: Ex-Researcher Mocks Agents Snapshatting Full Chat History on Safety Refusals(2 posts)→

Original post →

More from Safety

Safety channel →