LLMs enable at-scale deanonymization: up to 68% recall re-identifying HN and Reddit users

rickasaurus · x · 2026-09-25

A new arXiv paper by Simon Lermen, Nicholas Carlini, Florian Tramèr and colleagues shows LLM agents can deanonymize users at scale: with internet access, an agent re-identifies Hacker News and Anthropic Interviewer participants from pseudonymous profiles alone, matching hours of human investigator work. The pipeline uses LLMs to extract identity features, retrieve candidates via semantic embeddings, and verify matches. Unlike classic deanonymization work (e.g., Netflix Prize), it works on raw text across platforms. On three ground-truth datasets (HN–LinkedIn, Reddit movie communities, split Reddit histories) it substantially beats classical baselines, reaching up to 68% recall at 90% precision.

Related event: AI Pipeline Can Deanonymize Users for Just $2, Paper Shows(2 posts)→

Original post →

More from Safety

Safety channel →