Azure PII removal tool fails to protect 74% of info in MedQA, privacy paper finds
niloofar_mire · x · 2026-09-09
A new arXiv paper, "A False Sense of Privacy," shows synthetic data and paraphrasing are trivially re-identifiable — calling paraphrase "privacy preserving" is misguided.
Key findings:
- The authors propose a framework measuring real privacy risk via re-identification attacks, beyond checking explicit identifiers
- Seemingly innocuous auxiliary info (e.g. routine social activities) can infer sensitive attributes like age or substance use from sanitized text
- Azure's commercial PII removal tool fails to protect 74% of information in MedQA
- Differential privacy mitigates some risk but sharply reduces downstream utility
Conclusion: current sanitization gives a false sense of privacy; methods robust to semantic-level leakage are needed.
More from Safety
- Aidan Gomez: labs train on rewritten user data even under ZDR promises — josh_wills · 2026-09-09
- Quintin Pope: state-backed cybercriminals misusing AI pose a bigger threat than rogue labs — QuintinPope5 · 2026-09-09
- OpenAI Images V2.5 jailbroken, guardrails bypassed — flowersslop · 2026-09-09
- Researchers warn: don't feed commercially sensitive data to cloud LLMs as labs eye drug discovery — SumitGup · 2026-09-09
- Trigger a safety refusal, your whole history gets snapshotted: ex-Anthropic researcher mocks liability retention — suchenzang · 2026-09-09
- Anthropic pulls all ZDR options from Fable 5+ models citing safety — niloofar_mire · 2026-09-09