Trigger a safety refusal, your whole history gets snapshotted: ex-Anthropic researcher mocks liability retention
suchenzang · x · 2026-09-09
Former Anthropic researcher Suchenzang sarcastically flags a privacy wrinkle in new agent safety designs: trip a safety refusal and your entire conversation history may be snapshotted for "liability retention," to be combed through by an overenthusiastic AI agent while it eval-hacks "with nefarious intentions."
The underlying point is the tension between vendor safety and privacy terms: safety blocks now come bundled with long-term retention and review of user data — simultaneously absurd and unsettling.
More from Safety
- Provably private inference services promise prompts unreadable to providers — corbtt · 2026-09-09
- Google says EU DMA rules caused the biggest Search quality drop in its 29-year history — rickasaurus · 2026-09-09
- Developer questions whether OpenAI's ZDR policy is actually enforced — niloofar_mire · 2026-09-09
- Aidan Gomez: labs train on rewritten user data even under ZDR promises — josh_wills · 2026-09-09
- Quintin Pope: state-backed cybercriminals misusing AI pose a bigger threat than rogue labs — QuintinPope5 · 2026-09-09
- OpenAI Images V2.5 jailbroken, guardrails bypassed — flowersslop · 2026-09-09