Multimodal ASV Breaks Speaker Anonymization: EER Drops 15% with 5 Utterances

JohnsHopkins · hf · 2026-07-27

A study from Johns Hopkins University reveals privacy vulnerabilities in current speaker anonymization techniques. While many systems attempt to mask vocal characteristics, attackers can exploit multi-utterance, multimodal settings by combining acoustic, prosodic, and linguistic cues.

The research found that frame-level multimodal aggregation yields the best attack performance. Even with just five anonymized utterances, combining audio and text reduces the Equal Error Rate (EER) by over 15% relative to audio-only aggregation. This demonstrates that anonymization methods targeting only vocal traits leave substantial speaker-discriminative information accessible.

Original post →

More from Safety

Safety channel →