LLMs enable at-scale deanonymization: up to 68% recall re-identifying HN and Reddit users
rickasaurus · x · 2026-09-25
A new arXiv paper by Simon Lermen, Nicholas Carlini, Florian Tramèr and colleagues shows LLM agents can deanonymize users at scale: with internet access, an agent re-identifies Hacker News and Anthropic Interviewer participants from pseudonymous profiles alone, matching hours of human investigator work. The pipeline uses LLMs to extract identity features, retrieve candidates via semantic embeddings, and verify matches. Unlike classic deanonymization work (e.g., Netflix Prize), it works on raw text across platforms. On three ground-truth datasets (HN–LinkedIn, Reddit movie communities, split Reddit histories) it substantially beats classical baselines, reaching up to 68% recall at 90% precision.
Related event: AI Pipeline Can Deanonymize Users for Just $2, Paper Shows(2 posts)→
More from Safety
- Nat Lambert: Reid Hoffman's Framing Makes Open Source Look Far More Dangerous — natolambert · 2026-09-25
- One Neuron Is Enough to Bypass LLM Safety Alignment, NeurIPS 2026 Paper Shows — jonasgeiping · 2026-09-25
- Sen. Kelly proposes taxing top AI beneficiaries to fund workers, sparking pushback — robleclerc · 2026-09-25
- AI Model Muse Now Solves Captchas, Exposing Password Reset Security Flaw — illscience · 2026-09-25
- Irregular admits AI eval incidents were environment flaws, not rogue AI behavior — robleclerc · 2026-09-25
- Polymarket puts 16% odds on Anthropic announcing a full AI training pause this year — Polymarket · 2026-09-25